Not all cells in a cancer patient's tissue are equal. Even within a tumor, different regions vary dramatically in their gene activity -- some driving aggressive growth, others behaving more passively. Understanding which specific populations of cells are responsible for poor patient outcomes is an important goal in cancer research, because targeting those populations could lead to more precise and effective treatments.
Two technologies -- single-cell RNA sequencing (scRNA-seq) and spatial transcriptomics (ST) -- have made it possible to measure gene activity in individual cells and in precise tissue locations, respectively. However, these high-resolution technologies generate enormous amounts of data, and connecting the molecular patterns they reveal to actual patient outcomes (like survival or recurrence) remains technically challenging.
Existing methods can classify cells as tumor-like or normal based on gene expression, but linking these classifications to clinical outcomes -- does having more of these high-risk cells predict shorter survival? -- requires integrating genomic data with large patient outcome databases. This integration is the problem that DEGAS was developed to solve.
A particularly interesting biological question is whether tissue that looks normal under a microscope but shows abnormal gene activity patterns might serve as an early warning sign of cancer risk. Identifying such transcriptional precursors -- molecular changes that precede visible tissue changes -- could allow clinicians to identify high-risk patients before cancer is morphologically detectable.
DEGAS (Deep transfer learning for Gene Activity Signatures) is a computational framework that uses deep transfer learning to bridge two very different types of data: bulk tumor RNA sequencing from large patient cohorts with known clinical outcomes, and high-resolution single-cell or spatial transcriptomics data from tissue samples.
The approach works in two stages. First, DEGAS trains a deep learning model on bulk RNA sequencing data from the TCGA (The Cancer Genome Atlas) prostate cancer cohort, learning which patterns of gene expression across thousands of patients are associated with longer or shorter survival. This gives the model a survival-risk prediction capability grounded in large-scale clinical outcome data.
Second, DEGAS applies -- or transfers -- this learned survival association onto spatial transcriptomics slides, mapping the risk predictions onto individual tissue regions. The model identifies which specific physical locations in a tissue section show gene activity patterns most similar to the high-risk tumors it was trained on. This allows researchers to pinpoint which regions of a tissue slide are most associated with poor outcomes.
The method was validated by comparing it against Scissor, an established benchmark for linking single-cell data to clinical phenotypes. DEGAS showed that the number of high-risk regions identified in prostate cancer tissue slides increased predictably with disease stage -- a biologically expected result that confirms the method is capturing meaningful clinical signal rather than noise.
One of the most striking findings from applying DEGAS was the discovery of cancer-associated transcriptional signatures in tissue that appeared completely normal under a standard microscope. In both histologically normal prostate tissue and in adjacent-normal tissue from pancreatic cancer patients, DEGAS identified regions with gene activity patterns resembling those of high-risk tumors.
The genes enriched in these apparently normal but molecularly high-risk regions were associated with biological processes linked to cancer development: growth regulation and programmed cell death (apoptosis), inflammation, immune signaling, and autophagy (the cellular recycling process that can either suppress or promote cancer depending on context).
The researchers propose that these regions may represent transcriptional precursors to intraepithelial neoplasia -- the well-recognized premalignant state in which cells begin acquiring cancer-like properties before morphological changes become visible to a pathologist. If this interpretation is correct, DEGAS could detect molecular warning signs of cancer development before a biopsy would look abnormal.
A second clinical implication concerns cancer recurrence. Histologically normal tissue at the surgical margins around a tumor is sometimes left behind after cancer removal. If DEGAS can identify which of these margin regions carry hidden high-risk signatures, it could help explain why some patients whose surgery appeared complete still experience recurrence -- and potentially guide more targeted resection decisions.
The ability to detect molecular cancer risk in tissue that appears normal under a microscope has profound implications for clinical oncology. Current cancer diagnosis and staging rely heavily on what pathologists can see -- the morphological changes in cells and tissue architecture that indicate malignancy. DEGAS adds a molecular layer to this assessment that could identify risk before visible changes appear.
For prostate cancer specifically, where pathologists already use Gleason grading to assess tissue architecture, a complementary molecular risk map from a tool like DEGAS could help explain why cancers of the same Gleason grade sometimes behave very differently. If two tumors look identical under the microscope but show different transcriptional risk profiles in their surrounding normal tissue, the one with higher molecular risk in adjacent tissue might be more likely to recur after treatment.
The discovery of similar cancer-like signatures in adjacent-normal pancreatic tissue is particularly significant because pancreatic cancer is notoriously difficult to detect early. If DEGAS can identify transcriptional precursors in tissue adjacent to pancreatic tumors, it raises the possibility of monitoring such patients more aggressively for recurrence or new tumor development at these sites.
These findings support a broader principle in cancer biology: the field effect theory, which holds that cancer does not arise from a single cell in isolation but from a broader region of tissue under cellular stress, where many cells undergo subtle molecular changes before one eventually becomes overtly malignant. DEGAS provides a tool to measure this field effect computationally from transcriptomic data.
DEGAS represents an advance in the emerging field of computational spatial oncology -- the use of machine learning to extract clinically relevant information from the spatial organization of gene expression in tumor tissue. By connecting large-scale clinical outcome databases with high-resolution tissue-level molecular measurements, it bridges two types of data that have historically been analyzed separately.
The validation against a benchmark method (Scissor) and the biologically coherent result showing risk increasing with disease stage provide initial evidence that DEGAS captures genuine biological signal. However, as a conference abstract, the full methodological details, sample sizes, and statistical rigor of this work remain to be published in a complete peer-reviewed paper.
If validated in larger cohorts, DEGAS-style approaches could eventually be integrated into pathology workflows to provide molecular risk assessments alongside conventional histopathological reading -- adding a computational risk layer to standard biopsies that could guide treatment escalation or de-escalation decisions.
The generalizability of the method across two cancer types (prostate and pancreatic cancer) is encouraging and suggests that the transfer learning framework is not specific to one cancer biology. Future work applying DEGAS to other cancer types where spatial transcriptomics data are available could further establish it as a general tool for connecting molecular tumor heterogeneity to clinical outcomes.