Intratumoral Resolution of Driver Gene Mutation Heterogeneity in Renal Cancer Using Deep Learning.

Cancer Res 2022 AI 7 Explanations View Original
Original Paper (PDF)

Unable to display PDF. Download it here or view on PMC.

Plain-English Explanations
Pages 2-3
Intratumoral Heterogeneity: The Core Challenge

Intratumoral heterogeneity (ITH) means that different regions of the same tumor carry different genetic mutations. This creates a major clinical problem: a biopsy samples only one area, potentially missing driver mutations present elsewhere in the tumor that could influence prognosis and treatment resistance.

Multi-region sequencing studies have shown that to capture 75 percent of driver variants in clear cell RCC, an average of seven randomly sampled regions must be sequenced per tumor. This is impractical in clinical settings. Bulk sequencing of a single sample routinely misses sub-clonal events that may determine treatment outcomes.

Clear cell RCC (ccRCC) is a paradigm for ITH. Its three most frequently mutated driver genes, BAP1, PBRM1, and SETD2, can each be lost in only a subset of tumor cells, making slide-level gene status heterogeneous. Understanding where in a tumor these losses occur could guide smarter tissue sampling and predict clinical behavior more precisely.

Hematoxylin and eosin (H&E) stained slides are generated for every surgical specimen as a standard diagnostic step. If deep learning could infer genetic heterogeneity from these ubiquitous slides, it would create a low-cost, spatially resolved molecular profiling tool with no additional testing required.

TL;DR: Intratumoral genetic heterogeneity in ccRCC means that single-site biopsies miss critical driver mutations, and capturing the full mutational landscape currently requires impractical multi-region sequencing of seven or more samples per tumor.
Pages 4-8
A Multi-Cohort Study with IHC as Ground Truth

The study used immunohistochemistry (IHC) rather than bulk sequencing as ground truth for gene status. Validated IHC antibodies against BAP1, PBRM1, and H3K36me3 (a SETD2 readout) have positive and negative predictive values above 98 percent and provide spatial, cell-level resolution that sequencing cannot match. IHC staining of serial sections adjacent to H&E slides allowed direct spatial comparison.

Five cohorts were used spanning different imaging and sample types. The primary training cohort (WSI) included 1,282 whole slide H&E images from Mayo Clinic ccRCC patients. External validation used TCGA (363 patients with matched sequencing data) and two tissue microarray (TMA) cohorts from UT Southwestern with 118 and 365 patients respectively. A patient-derived xenograft (PDX) TMA cohort of 46 patients was used to test whether models trained on human tissue generalize to animal models.

Two distinct model types were trained for each gene. A slide-level model predicted whether gene loss was present anywhere on the slide, paralleling bulk sequencing approaches. A region-level model predicted gene status within 1mm-diameter local regions, enabling spatial mapping of intratumoral heterogeneity. This two-scale design directly addressed the clinical problem of sub-clonal spatial resolution.

The slide model used a multiple instance learning (MIL) architecture with an attention mechanism that learned to weight patches more heavily if they appeared loss-like. The region model used a fully convolutional VGG19 network generating continuous activation maps across the tumor. Stain normalization was applied to TMA cohorts to reduce inter-institutional color variation.

TL;DR: IHC-based ground truth at cellular resolution enabled training of both slide-level and region-level deep learning models for three driver genes across five patient cohorts spanning different institutions and sample formats.
Pages 11-13
BAP1 Loss Is Strongly Linked to Tissue Morphology

At the whole-slide level, all three genes showed significant AUCs on the held-out WSI test set: BAP1 = 0.87, PBRM1 = 0.77, SETD2 = 0.71. BAP1 showed the strongest relationship between morphology and gene loss, consistent with its known impact on nuclear grade and tissue architecture.

External validation on TCGA confirmed these results with AUCs of 0.77, 0.69, and 0.60 for BAP1, PBRM1, and SETD2. These values compare favorably to published pan-cancer models that trained and tested on TCGA directly, where the BAP1 AUC was only 0.65. The performance advantage is likely due to the larger BAP1 loss training set and superior IHC-based ground truth.

Negative predictive values were particularly high: the model correctly ruled out BAP1 loss in 94 percent of predicted-WT slides in WSI testing and 93 percent in TCGA. By examining the top 60 percent of samples ranked as most likely BAP1 loss, 96 percent of true BAP1 loss cases were captured. High sensitivity with a clean high-confidence WT zone makes this practically useful for ruling out loss.

TL;DR: BAP1 loss produced the strongest morphological signal with AUC 0.87 on held-out training data and 0.77 on independent TCGA slides, with a negative predictive value of 94 percent enabling reliable rule-out of BAP1 loss.
Pages 13-15
Spatial Resolution of Intratumoral Heterogeneity

The region-level BAP1 model achieved AUC 0.89 on WSI test regions, essentially matching its slide-level counterpart. This means the model can not only detect that BAP1 loss exists somewhere on a slide, but can localize where within the tumor it occurs at 1mm spatial resolution.

On independent TMA cohorts, the region model achieved AUC 0.77 on TMA1 (118 patients) and AUC 0.84 on TMA2 (365 patients). On the PDX cohort, where human tumor stroma was replaced by mouse tissue, the model still achieved AUC 0.80. This PDX result demonstrates that the morphological features associated with BAP1 loss are intrinsic to tumor cells, not driven by interactions with human stroma.

In localized-loss cases, where only part of the tumor shows BAP1 loss, the model correctly identified the most likely loss area better than 98.7 percent of random area selections. This means that if targeted sequencing were guided by the model, it would reliably direct sampling toward the regions harboring BAP1 mutations, turning an impossible problem (sampling all 7+ required regions) into a tractable guided approach.

TL;DR: The region-level model resolves BAP1 loss within individual tumors at 1mm resolution, outperforming random biopsy site selection by more than 98 percent and maintaining performance in independent cohorts including patient-derived xenograft tissue.
Pages 15-16
Gene-Specific Morphological Signatures

To understand what features drive BAP1 predictions, the team analyzed 36 nuclear morphology features extracted from all tumor nuclei across the WSI cohort. Random forest classifiers based on these nuclear features alone achieved AUC 0.72 for BAP1, confirming that nuclear morphology is a genuine contributor to the deep learning signal.

BAP1 mutant tumors consistently showed larger nuclei and lighter nuclear hematoxylin staining (open chromatin appearance) compared to wildtype tumors. These findings align with the known role of BAP1 in chromatin regulation and its established link to higher nuclear grade in ccRCC.

Importantly, BAP1 and PBRM1 predictions were negatively correlated (Spearman coefficient -0.33), consistent with the biological observation that BAP1 and PBRM1 mutations are largely mutually exclusive in ccRCC. PBRM1 and SETD2 showed positive correlation in predictions, mirroring their known co-occurrence. This means the models independently rediscovered established genetic co-exclusivity from morphology alone.

TL;DR: BAP1 loss is marked morphologically by larger, lighter nuclei consistent with its chromatin regulatory role, and model predictions for different genes recapitulate the known mutual exclusivity of BAP1 and PBRM1 mutations.
Pages 16-17
BAP1 Predictions Correlate with Patient Survival

Cox regression in TMA2 (365 patients) showed that the BAP1 model's continuous prediction score significantly predicted disease-specific survival with a hazard ratio of 5.43 and a concordance index of 0.65. This performance exceeded stratification based on IHC-measured BAP1 status, which yielded a c-index of only 0.55.

Kaplan-Meier survival analysis stratified by model predictions showed an improved hazard ratio (2.53 versus 1.8) and p-value (0.003 versus 0.11) compared to direct IHC-based BAP1 classification. This suggests the model captures survival-relevant morphological information beyond what IHC protein expression conveys.

These findings raise the intriguing possibility that the morphological state identified by the deep learning model represents a broader cellular phenotype of which BAP1 loss is one manifestation. Other mechanisms could produce the same high-risk morphological signature, potentially explaining why the model outperforms the direct molecular test in survival prediction.

TL;DR: The BAP1 deep learning score predicted kidney cancer-specific survival better than IHC-based BAP1 status measurement, suggesting the model captures a broader high-risk morphological phenotype with direct prognostic relevance.
Pages 18-19
Linking Morphology to Genetics at Sub-Clonal Resolution

This study established that deep learning models trained on H&E images can resolve intratumoral genetic heterogeneity at 1mm spatial resolution, a capability that sequencing-based approaches cannot practically achieve. The approach converts the ubiquitous standard pathology slide into a spatially resolved molecular profiling tool.

While the models are not yet ready to replace IHC or sequencing in clinical practice, they demonstrate proof of concept for a low-cost in-silico method that could guide targeted sequencing, stratify patients for clinical trials, or serve as a triage tool to prioritize molecular testing for BAP1 status.

The preservation of model performance in patient-derived xenograft tissue and across multiple independent human cohorts strongly supports the generalizability and biological grounding of these predictions, positioning this approach for prospective validation and eventual clinical translation in kidney cancer diagnosis and risk stratification.

TL;DR: Deep learning on H&E slides can spatially resolve intratumoral genetic heterogeneity in ccRCC, providing a clinically practical low-cost alternative to multi-region sequencing for detecting driver gene loss and predicting patient survival.
Citation: Open Access, 2022. Available at: PMC9373732.