Prostate cancer (PCa) is the second most common cancer and the fifth leading cause of cancer-related deaths in men worldwide. While many patients undergo treatments like radical prostatectomy or radiation therapy, a significant number experience biochemical recurrence (BCR) -- a rise in prostate-specific antigen (PSA) levels after treatment that often signals cancer returning locally or spreading to distant sites.
BCR is an important warning sign: it reduces survival and requires prompt treatment decisions. Identifying which patients are at highest risk of BCR early allows clinicians to initiate secondary therapies quickly, potentially delaying disease progression and improving long-term outcomes.
Current tools for predicting BCR include traditional nomograms like the CAPRA score, which combine clinical variables to estimate recurrence risk. However, these tools rely on single or limited parameters, may be subject to reader variability, and often fail to capture the full complexity of prostate cancer's biological heterogeneity.
Machine learning (ML) offers an alternative approach that can integrate large amounts of complex, multi-dimensional data -- including clinical, imaging, and pathological information -- to potentially predict BCR more accurately than traditional methods. This meta-analysis systematically evaluated how well ML models perform for this task.
Following PRISMA guidelines, researchers searched four major databases -- PubMed, Web of Science, Embase, and Cochrane -- using the terms 'prostate cancer,' 'machine learning,' and 'biochemical recurrence.' The search also captured related terms such as PSA recurrence and biochemical failure.
From an initial pool of 1,095 articles, 16 studies were ultimately included after rigorous screening. Studies focused on radiomics alone, those predicting non-BCR outcomes, or those without reported AUC (area under the curve) data for both ML and a traditional comparison model were excluded.
The 16 included studies covered approximately 17,316 prostate cancer patients from multiple countries and institutions, spanning publications from 2007 to 2024. All were prospective or retrospective cohort studies evaluating ML model performance against traditional clinical models.
Since AUC -- a combined measure of sensitivity and specificity -- was reported in all 16 studies, it was chosen as the primary performance metric. A random-effects model was used for the meta-analysis to account for moderate heterogeneity between studies (I2 = 49.5%), and subgroup analyses examined model type, data type, and prediction time interval.
The meta-analysis revealed a pooled AUC of 0.82 (95% CI: 0.81-0.84) across all 82 ML models included from the 16 studies. This result indicates reliable discriminative ability -- meaning these models can meaningfully distinguish between patients who will and will not experience BCR after treatment.
This is notably superior to the CAPRA nomogram, a widely used traditional tool, which achieved a pooled AUC of only 0.73 in a comparable meta-analysis. The difference underscores ML's advantage in processing complex, multi-dimensional clinical data that traditional statistical models cannot fully leverage.
Among the 82 models evaluated, the majority (76.8%) were non-deep learning models such as logistic regression and random forest. Deep learning models (18.3%) and hybrid ML + deep learning models (4.9%) were less common but showed consistently strong performance, with pooled AUCs of 0.83 each -- slightly outperforming traditional ML models (AUC 0.81).
Both logistic regression and random forest -- the two most commonly used model types -- achieved comparable pooled AUCs of 0.84 each, demonstrating that even established, interpretable ML approaches can perform strongly for BCR prediction. The stability of results across model types lends confidence to the overall findings.
A key finding was that imaging data significantly enhanced model performance. Models incorporating imaging data achieved a pooled AUC of 0.82, compared to 0.78 for models using only clinical or pathological data. This confirms that information from scans such as MRI -- including tumor size, location, and tissue characteristics -- provides unique predictive value beyond standard clinical variables.
The majority of models (85.4%) used hybrid data combining imaging, pathological, and clinical laboratory data. This multimodal approach showed more stable performance than single-data models, suggesting that integrating different types of patient information leads to more comprehensive and robust predictions.
A standout example was a study by Hou et al., which combined MRI data with AI to predict BCR non-invasively. Their AI model outperformed traditional CAPRA and CAPRA-S nomograms across 1-year, 2-year, and 3-year prediction intervals, demonstrating that combining rich imaging data with ML can reduce the need for invasive biopsies.
Shiradkar et al. developed a multimodal model using convolutional neural networks (CNN) and imaging features, achieving an AUC of 0.860 -- substantially higher than CAPRA (0.684) and CAPRA-S (0.705). This example illustrates the particular power of deep learning approaches when applied to complex multi-source data.
ML models showed varying performance depending on how far into the future they were predicting BCR. The 1-year BCR prediction was the most accurate, achieving a pooled AUC of 0.86 -- the highest across all time intervals examined. This reflects that models can most reliably identify patients at imminent risk of recurrence.
Predictive accuracy decreased modestly for the 2-year (AUC 0.82) and 3-year (AUC 0.80) intervals, consistent with the idea that predicting further into the future introduces more uncertainty as cancer biology evolves and new variables emerge. However, all time-interval estimates remained above 0.80 -- a threshold generally considered strong.
Interestingly, the 5-year prediction recovered to an AUC of 0.82, comparable to the 1-year estimate. This suggests that ML models may be able to identify long-term risk patterns in patient data -- such as aggressive tumor biology or underlying genetic factors -- that remain detectable even years after initial treatment.
Notably, Tan et al. achieved AUCs of 0.86, 0.86, and 0.88 at 1-year, 3-year, and 5-year intervals respectively, outperforming traditional nomograms at all time points. These results support ML as a reliable tool not just for near-term, but also for sustained long-term recurrence surveillance.
Across multiple studies, ML models consistently outperformed traditional clinical nomograms. The CAPRA scoring system, one of the most widely used tools for BCR prediction, achieved a pooled AUC of 0.73 in a prior meta-analysis -- significantly below the 0.82 pooled AUC for ML models in this study. This 9-point gap in AUC represents a meaningful clinical difference in the ability to correctly risk-stratify patients.
The advantage of ML stems from its ability to handle high-dimensional, nonlinear data. Traditional nomograms rely on a fixed set of pre-selected variables and linear relationships, while ML algorithms can automatically identify complex interactions among dozens or hundreds of clinical, imaging, and pathological features simultaneously.
Deep neural networks (DNN) showed particular promise: one study found that DNN achieved an AUC of 0.84 for 3-year BCR prediction, outperforming CAPRA score, logistic regression, random forest, k-nearest neighbors, and Cox regression in the same patient cohort. These results suggest that the most sophisticated ML architectures may be especially valuable when complex patterns need to be detected.
Despite these advantages, the study acknowledges that most evidence comes from retrospective data, and heterogeneity across studies -- including differences in patient populations, treatment modalities, and data sources -- may limit direct comparisons. Prospective, multicenter clinical validation remains necessary before widespread adoption.
The findings have direct implications for clinical practice. By identifying patients at high risk of BCR earlier and more accurately, ML tools could guide clinicians toward timely decisions about salvage radiation, hormone therapy, or clinical trial enrollment -- potentially improving survival outcomes.
Non-invasive prediction approaches that combine MRI imaging with AI are particularly promising because they can provide accurate BCR risk estimates without requiring additional biopsies. This reduces patient discomfort, infection risk, and healthcare costs while maintaining high predictive accuracy.
For low- and middle-income countries, ML models offer the potential to integrate existing clinical data into cost-effective prediction tools that could reduce reliance on expensive specialized tests. However, challenges remain around data quality, algorithm transparency, and the need for clinician training before broad implementation.
Future priorities include developing and validating multimodal models that incorporate newer data sources such as PSMA-PET/CT imaging, and conducting large-scale multicenter prospective trials to confirm ML's clinical utility across diverse healthcare settings. The authors emphasize that AI should complement -- not replace -- clinical expertise.