Early detection of prostate cancer could significantly reduce mortality, but current biomarkers are unreliable. PSA (prostate-specific antigen), the most widely used screening test, has poor specificity and sensitivity, leading to many unnecessary biopsies or missed diagnoses. More sophisticated biomarkers like the Prostate Health Index (measuring three PSA forms) achieve AUCs of only 0.70-0.77, still leaving substantial diagnostic uncertainty.
A fundamental challenge in biomarker discovery is that prostate tumors are extraordinarily complex. Each tumor involves altered interactions between thousands of genes in multiple cell types, and these interactions vary between patients with the same diagnosis and even between different regions of the same tumor. Data-driven approaches that ignore this complexity tend to discover biomarkers that do not replicate across clinical settings.
Urine has emerged as a promising sample type because it captures signals from prostate tissue more directly than blood. Recent proteomic analyses of urine identified 200+ proteins that collectively reached AUCs of 0.71-0.81 for prostate cancer diagnosis. Combining urine biomarkers with clinical factors or imaging pushed AUCs up to 0.89. However, these studies were still limited by bulk analyses that miss signals from individual cell types and by data-driven feature selection that lacks mechanistic grounding.
This study hypothesized that more reliable biomarkers could be found by leveraging a fundamental biological principle: within a single tumor, cells exist at different stages of malignant transformation. By mapping these stages using spatial transcriptomics and modeling the transformation trajectory using pseudotime, the genes most critical to malignant transformation could be systematically identified and validated as measurable biomarkers.
Spatial transcriptomics (ST) is a technology that measures gene expression at thousands of individual spatial locations (spots) across a tissue section while preserving information about where each measurement comes from. Unlike traditional bulk RNA sequencing, which produces a single average across the whole sample, ST reveals how gene activity varies across different regions of the tumor and adjacent normal tissue.
The researchers retrieved ST data from the publicly available HEST-1k repository, which contains more than 1,000 spatial transcriptomics profiles. After filtering for human prostate cancer samples with sufficient spot coverage and excluding technical replicates, they worked with 12 high-quality ST samples from six patients. A pathologist manually annotated each spot as representing normal gland, prostatic intraepithelial neoplasia (PIN), or cancer, with cancer spots further classified into five Gleason grade groups representing increasing malignancy.
The final analysis included 19,695 spots representing an average of 1,641 spots per sample, with each spot measuring approximately 4,800 unique genes. This resolution captured the spatial heterogeneity of tumor progression: some samples were predominantly benign with only small malignant foci, while others contained extensive cancer regions, allowing the full spectrum of malignant transformation to be analyzed.
The distribution of different cancer grades varied substantially between samples, reflecting the known heterogeneity of prostate tumors and reinforcing the need for an approach that could find reliable signals despite this complexity. Having both benign and malignant cells in the same samples was essential for the next step: modeling the trajectory of malignant transformation using pseudotime.
Pseudotime (PT) is a computational method that organizes cells (or spatial spots in this case) along a developmental trajectory based on the similarity of their gene expression profiles. Each spot is assigned a value from 0 to 1, where values near 0 indicate expression profiles similar to normal tissue and values near 1 indicate profiles most divergent from normal, representing the most advanced malignant transformation.
The pseudotime calculation uses diffusion maps, a mathematical technique that captures non-linear relationships in gene expression data and is more robust to noise than simpler approaches like PCA. A root spot representing benign tissue is identified, and the algorithm then calculates the probability of transitioning between each pair of spots based on expression similarity, tracing the path from normal to malignant states.
When pseudotime was calculated and compared against the histological grades annotated by the pathologist, PT values correlated significantly with cancer grade (p < 0.001 by Mann-Whitney U test). Normal gland spots had the lowest PT, while higher-grade cancer spots had progressively higher PT values. This validated the central hypothesis that PT models malignant transformation.
Copy number aberration (CNA) analysis, which detects regions of the genome that are abnormally duplicated or deleted in cancer cells, further confirmed the PT model. Spots with higher PT values showed more extensive CNAs, consistent with the accumulation of genomic instability during cancer progression. The convergence of histological grade, PT, and CNA provided a three-way validation of the pseudotime model.
Using pseudotime across all 12 spatial transcriptomics samples, the researchers identified genes whose expression was most strongly correlated with pseudotime, meaning genes whose activity systematically changed as tissue transformed from normal to malignant. Genes positively correlated with PT increase as cancer progresses; those negatively correlated decrease.
The gene set identified included known prostate cancer genes like KLK3 (PSA), AMACR, and KLK2, validating that the approach recovered biologically relevant cancer genes. Novel candidates like SPON2, TMEFF2, STEAP2, and CRISP3 also emerged as highly correlated, suggesting new biomarker opportunities. Functional enrichment analysis showed these genes are associated with pathways including androgen response, epithelial-mesenchymal transition, interferon response, and apoptosis.
Drug target analysis linked these genes to potential therapeutic agents including pembrolizumab, nivolumab, ipilimumab (immunotherapy drugs), and platinum-based chemotherapy, suggesting that genes most strongly associated with malignant transformation are also clinically actionable targets.
Immunohistochemistry validation of the protein SPON2 in prostate tissue confirmed increasing expression with cancer grade, providing physical evidence that the computationally identified biomarkers reflect real changes in prostate tissue protein abundance across the spectrum of malignancy.
The candidate biomarker genes were validated across multiple independent data types from more than 2,000 patients: bulk RNA expression from prostate tissue, single-cell RNA sequencing, serum proteomics from the UK Biobank (2,923 proteins from 500,000+ participants), urine proteomics, and immunohistochemistry. This multi-layered validation across different biological materials and patient populations was designed to test whether the computationally identified genes represent truly robust signals rather than dataset-specific artifacts.
Machine learning prediction models (ridge regression) were built using the candidate biomarker proteins from urine and blood. The urine-based model achieved an AUC of 0.92 for prostate cancer diagnosis, substantially outperforming the serum PSA model and the UK Biobank serum-based biomarker model (AUC 0.69). This large performance gap between urine and blood reflects the proximity of urine to prostate tissue and the richer cancer-associated protein signal in post-digital rectal exam urine samples.
Beyond diagnosis, the biomarkers also predicted cancer grade, with higher pseudotime gene expression correlating with more aggressive Gleason scores. This is clinically significant because it suggests these markers could help distinguish low-grade, indolent cancer (which may not require treatment) from high-grade, aggressive disease (which does), a distinction that PSA cannot make.
To confirm that the improvement was specifically due to the pseudotime-guided gene selection rather than chance, the same prediction pipeline was repeated 1,000 times with randomly selected genes. The randomly selected genes consistently achieved lower AUCs, confirming that the pseudotime-guided biomarkers provide biologically meaningful predictive power beyond what would be expected at random.
The key methodological innovation of this study is the combination of three technologies that individually have been applied to cancer research but have not previously been integrated: spatial transcriptomics to capture the geography of tumor heterogeneity, pseudotime to model the trajectory of malignant transformation, and machine learning to translate molecular signals into clinically measurable predictions.
The use of pseudotime to prioritize biomarkers has an important advantage over purely data-driven approaches: it anchors biomarker selection to the biology of cancer development. Genes that are most strongly associated with the pseudotime trajectory represent molecular changes that consistently distinguish benign from malignant tissue across patients and tumor regions, which should make them more likely to replicate across different clinical settings.
The authors explicitly propose this framework as generalizable to other cancer types. The same pipeline, applied to spatial transcriptomics data from other cancers where both benign and malignant cells exist in the same tumor, could identify cancer type-specific biomarker sets. The code and data for this study were made freely available to facilitate such extensions.
Limitations include the small number of spatial transcriptomics samples (12 samples from 6 patients) and the retrospective nature of the validation datasets. Prospective clinical studies testing whether these urine biomarkers can guide real-world screening decisions are needed before the approach can influence clinical practice.
This study demonstrates that integrating spatial transcriptomics, pseudotime analysis, and machine learning can identify prostate cancer biomarkers with clinical-grade diagnostic accuracy. An AUC of 0.92 from urine proteins is among the highest reported for prostate cancer diagnosis with a non-invasive test, significantly exceeding the performance of PSA.
The fact that the biomarkers predicted cancer grade, not just cancer presence, addresses one of the central clinical challenges in prostate cancer management: the need to distinguish clinically significant cancers that require treatment from low-risk cancers that can be safely monitored. Grade-predictive biomarkers could reduce unnecessary treatment and its associated side effects.
The validation across multiple independent patient cohorts and sample types (tissue, blood, urine) provides strong evidence for robustness. The convergent signals across different biological compartments suggest these genes reflect fundamental features of prostate cancer biology rather than tissue-type-specific artifacts.
The next step is prospective validation: testing these urine biomarkers in clinical trials where men present for prostate cancer screening, to confirm whether the diagnostic accuracy observed in retrospective datasets translates to real-world screening performance and clinical utility.