Prostate cancer is the most commonly diagnosed solid cancer in men globally and ranks sixth among cancer-related deaths in men. When caught early, it is highly treatable -- but the diagnostic pathway is imperfect. PSA (prostate-specific antigen) testing has reduced prostate cancer mortality by over 50%, yet it also drives overdiagnosis: many detected cancers are indolent and would never have caused harm, leading to unnecessary anxiety and treatment.
Standard imaging with multiparametric MRI (mpMRI) has improved the detection of clinically significant disease by combining structural T2-weighted imaging with functional sequences that reveal how aggressively a tumor is behaving. Suspicious findings are reported using the PI-RADS scoring system (Prostate Imaging Reporting and Data System), which assigns lesions a score from 1 (very unlikely cancer) to 5 (highly likely cancer). Despite its value, mpMRI still suffers from moderate inter-reader variability and low specificity.
A second opinion -- having a different expert re-read the same scan -- is an established way to reduce diagnostic errors in medicine. But second opinions are expensive, slow, and often inaccessible, particularly in low- and middle-income countries. AI tools that automatically analyze MRI images could provide an always-available, consistent second opinion that does not depend on access to scarce radiology expertise.
This systematic review was conducted to evaluate the evidence on whether AI can meaningfully improve diagnostic accuracy for prostate cancer on MRI, and to assess whether current AI tools are mature enough to serve as reliable second-opinion systems in clinical practice.
The review followed PRISMA guidelines (Preferred Reporting Items for Systematic Reviews and Meta-Analyses), searching six major databases -- PubMed, Embase, Ovid MEDLINE, Web of Science, Cochrane Library, and IEEE Xplore -- for studies published between January 2019 and April 2024. The five-year window was chosen to capture the rapid recent growth of AI in oncology imaging.
Eligible studies had to involve AI analysis of prostate MRI for cancer diagnosis. The search started with 646 studies; after removing abstracts, non-English publications, duplicates, and studies that used MRI without AI, the final included set was seven studies. While small, this reflects the high inclusion criteria rigor -- studies without clear lesion annotation methodology or full-text methodology were excluded.
Study quality was assessed using the QUADAS-2 (Quality Assessment of Diagnostic Accuracy Studies) tool, which evaluates four domains: patient selection, index test (the AI), reference standard (histopathology), and flow and timing of assessments. Studies were rated as low, moderate, or high risk of bias in each domain. Most studies carried moderate to high risk overall, primarily because of their retrospective designs -- studies that look back at existing patient data rather than prospectively enrolling new patients.
The seven included studies used different AI approaches and focused on different aspects of prostate cancer detection. Four used deep learning, two used traditional machine learning (including random forest and support vector machine classifiers), and three specifically used convolutional neural networks (CNNs). One study -- Ström et al. -- was prospective; all others were retrospective. Performance ranged widely depending on the task and data quality.
Key individual study results included: Salman et al. achieved 97% accuracy on similar test images and 89% on diverse biopsy images using a CNN for classifying prostate tissue biopsy images. Ström et al. achieved a remarkable AUC of 0.997 for differentiating benign from malignant biopsy cores using a deep neural network. Hosseinzadeh et al. showed their deep learning model detected PI-RADS 4 or higher lesions with 87% sensitivity and Gleason score above 6 lesions with 85% sensitivity -- demonstrating useful localization capability.
Lee et al. directly compared two AI approaches: texture-based machine learning achieved AUCs of 0.854-0.861 with high specificity (0.71-0.78), while image-based deep learning achieved higher sensitivity (0.946) but lower specificity (0.643) and AUC (0.802). This trade-off illustrates a common pattern -- different AI architectures have different precision-recall profiles, and the best choice depends on whether the clinical priority is avoiding false negatives (missing cancer) or false positives (unnecessary biopsies). Yu et al. found their AI-based PI-RADS system outperformed more than 70% of ordinary radiologist readers.
Khosravi et al. combined biopsy report data with MRI images, achieving AUC 0.89 for distinguishing cancer from benign cases and 0.78 for separating high-risk from low-risk disease -- demonstrating that AI combining pathology and imaging outperforms imaging alone. Hectors et al. specifically addressed the challenging problem of PI-RADS 3 lesions (equivocal findings where diagnostic uncertainty is highest), achieving AUC 0.76 with a random forest classifier in this difficult subgroup.
The review frames AI's most practical near-term role as a second opinion tool rather than a replacement for radiologists. A second opinion is particularly valuable for PI-RADS 3 lesions -- the indeterminate middle group where radiologists disagree most. One referenced study of 950 cases found that when affiliated centers' PI-RADS 3 calls were reviewed by expert radiologists, 35.7% of lesions were downgraded and 15% were upgraded, illustrating how much interpretive variability exists at this threshold. AI could provide a consistent, instantaneous second interpretation that captures patterns human readers miss.
The review also cites evidence that even high-performing AI systems can catch cancers that experienced pathologists missed. The Paige Prostate AI system -- evaluated on 600 biopsies from 100 patients -- achieved 99% sensitivity and 93% specificity and detected patients that three experienced histopathologists had previously failed to diagnose. This type of safety-net performance is where AI adds unique value: finding rare patterns in large slide collections that human attention may overlook.
An important equity argument for AI is its potential to narrow healthcare access gaps. In low- and middle-income countries where experienced prostate MRI readers are scarce, an AI second-opinion system could provide consistent diagnostic support that would otherwise be unavailable. This democratization of expertise represents a different value proposition than efficiency gains in well-resourced centers.
The review's most significant limitation is the small number of qualifying studies (seven) despite searching six databases. This reflects strict inclusion criteria but also the genuine scarcity of rigorously designed AI prostate MRI studies with clear methodology, appropriate reference standards, and complete outcome reporting. Many published AI studies in this area remain exploratory proof-of-concept work rather than clinical validation studies.
Six of the seven included studies were retrospective -- designed after the fact using pre-existing patient data. This introduces selection bias and limits how confidently findings can predict performance in a live clinical setting where the AI's outputs might influence decisions differently than when they are compared retrospectively. The one prospective study (Ström et al.) achieved notably strong results, suggesting that well-designed prospective validation is feasible and necessary.
Heterogeneity across studies -- different AI algorithms, dataset sizes, MRI protocols, patient populations, and outcome definitions -- made direct comparison impossible. The review also noted that most AI models were trained and tested on single-center, single-population datasets, limiting generalizability. Whether models calibrated for one country's imaging protocols and demographics will perform equally in different health systems has not been established.
The systematic review concludes that AI has demonstrated meaningful potential for improving prostate cancer diagnosis from MRI -- achieving high AUC values, outperforming average radiologists in some comparisons, and showing particular promise as a consistent second opinion for equivocal findings. The evidence supports continued development but does not yet justify replacing radiologists or deploying AI without human oversight.
The field's next priorities are prospective multi-center validation studies that test AI systems on diverse patient populations across multiple institutions and imaging platforms. Current evidence is almost entirely from retrospective single-site work, and a tool that performs well in one academic center's dataset may perform differently in a community hospital or a different country's patient population.
The review also calls for research to expand AI's role beyond detection: AI could assist with staging, treatment selection, monitoring treatment response, and predicting cancer progression. These applications require different types of AI tools and validation evidence, but represent the direction needed to fully realize AI's potential as an integrated partner across the complete prostate cancer care pathway.