Non-muscle-invasive bladder cancer accounts for the majority of bladder tumors and carries a 70 to 80 percent recurrence rate after treatment. NMIBC includes papillary tumors confined to the bladder mucosa (Ta), invasion of the subepithelial connective tissue (T1), and carcinoma in situ (Tis). Approximately 50 percent of untreated NMIBC patients will progress to muscle-invasive disease, making early and accurate recurrence prediction critical for patient outcomes.
Bladder cancer is the tenth most common cancer globally and imposes substantial economic burden. In the United States, bladder cancer accounts for nearly 3.7 billion dollars in direct costs, with per-patient treatment costs for NMIBC ranging from 5,594 to 9,554 dollars. The global burden is rising as BCG therapy availability decreases, increasing pressure on predictive tools that might reduce unnecessary surveillance procedures.
Current surveillance relies on invasive cystoscopy, cytology, imaging, and biopsy performed repeatedly over time. Standard risk stratification uses the EORTC and CUETO scoring systems based on tumor stage, grade, size, multiplicity, recurrence rate, and presence of carcinoma in situ. These systems can underestimate recurrence in low-risk disease and overestimate it in high-risk disease, leaving substantial room for improvement.
This comprehensive review analyzed 70 studies from the past decade covering radiomics, histopathological, clinical, and genomic markers for NMIBC recurrence prediction. The review evaluates both individual and combined biomarker approaches, with a goal of identifying future AI-based directions that could enable personalized management and reduce dependence on repeat invasive procedures.
Radiomics markers extract quantitative imaging features from CT, MRI, and PET/CT scans to predict NMIBC recurrence without tissue sampling. Four studies evaluated radiomics-based approaches. DWI-MRI radiomics achieved the highest performance with accuracy of 0.915 to 0.926 across two studies using CNN and AlexNet architectures. FDG PET/CT achieved accuracy of 0.886 to 0.90 in two studies using SVM and U-Net classifiers.
Deep learning models applied to radiomics features consistently outperformed traditional machine learning classifiers. CNN-based analysis of DWI-MRI images extracted features related to tumor vascularity and cellularity to predict recurrence. AlexNet models applied to ADC maps from DWI sequences achieved accuracy of 0.915 with sensitivity of 0.87 and specificity of 0.95, demonstrating the potential of transfer learning for NMIBC surveillance.
FDG PET/CT radiomics provided prognostically meaningful features linked to tumor metabolic activity. SVM classifiers trained on standardized uptake value and texture features from PET/CT achieved sensitivity of 0.93 and specificity of 0.85. The non-invasive nature of radiomics makes these approaches attractive complements to cystoscopy, though external validation remains limited across all included radiomics studies.
A combined radiomics and clinical marker model outperformed either modality alone. Xu and colleagues combined 32 textural radiomics markers with muscle-invasive status as a clinical marker in 71 patients. The combined model achieved accuracy of 0.809 and AUC of 0.838 for 2-year recurrence prediction, compared to 0.755 accuracy for radiomics alone, demonstrating the complementary value of integrating imaging features with clinical information.
AI models applied to histopathological whole-slide images achieved up to 90 percent accuracy for NMIBC recurrence prediction. Seven studies evaluated histopathological markers. The most powerful approach used SVM trained on 79 nuclear morphology features extracted from pathologist-annotated regions, achieving accuracy of 0.90. Nuclear characteristics including size, shape, and distribution captured prognostically relevant information invisible to routine pathological review.
Squamous and glandular differentiation patterns in bladder tumors were identified as significant histopathological predictors of recurrence. These architectural features, detectable in standard hematoxylin and eosin stained sections, were associated with increased recurrence risk and provided independent predictive value beyond tumor grade and stage. AI quantification of these patterns offers reproducibility advantages over subjective pathological assessment.
Deep learning applied to whole-slide images combined with clinical markers improved 5-year recurrence prediction to AUC of 0.76. Lucas and colleagues used U-Net segmentation followed by VGG16 feature extraction and bidirectional GRU classification across 200 histopathological markers combined with clinical variables including tumor stage, intravesical chemotherapy, and smoking history. The combined model outperformed either histopathology or clinical data alone.
Tumor budding quantification from histopathological images added independent prognostic value for recurrence prediction. Automated detection of tumor budding at the invasive front captured features related to epithelial-mesenchymal transition and invasive potential. These image-derived tumor microenvironment features, combined with standard clinical parameters, separated patients into distinct recurrence risk groups with statistically significant survival differences.
Seventeen studies evaluated clinical and clinicopathological markers for NMIBC recurrence, with most using statistical analysis rather than advanced AI. The inflammatory marker neutrophil-to-lymphocyte ratio emerged as a consistent predictor, with NLR greater than 2.5 independently predicting recurrence. The modified Glasgow Prognostic Score, erythrocyte sedimentation rate, and metabolic markers including dopaquinone and leucine in urine also showed significant associations with recurrence risk.
Treatment modality was among the strongest clinical predictors of recurrence in comparative studies. A network meta-analysis of 12,464 patients across intravesical therapy types found gemcitabine achieved the highest recurrence prevention ranking with a surface under cumulative ranking curve of 0.92, followed by BCG at 0.82 and interferon at 0.78. Narrow-band imaging during TURBT reduced 1-year recurrence rates from 51 percent to 33 percent compared to standard white-light resection.
Restaging TURBT before BCG therapy significantly reduced recurrence rates in high-grade NMIBC. A study of 1,021 patients found that single TURBT was associated with a recurrence rate of 0.772, while restaging TURBT reduced this to 0.616. Female gender was independently associated with higher recurrence risk and impaired BCG response in a meta-analysis of 23,754 patients, highlighting the need for sex-stratified treatment planning.
Cystoscopy follow-up delay beyond 62 days significantly increased recurrence risk. A study of 407 patients found delays of 2 to 5 months and greater than 5 months progressively increased recurrence probability, demonstrating that adherence to surveillance schedules is itself a clinically modifiable recurrence risk factor. Smoking history showed poor association with recurrence in a 722-patient study, limited by small numbers of patients who actually quit smoking.
Twenty-five studies evaluated genomic biomarkers for NMIBC recurrence, with TERT and FGFR3 mutations being the most widely studied across six studies. TERT promoter mutation detection achieved accuracy of 0.933 and sensitivity of 1.00 in the highest-performing study, while the combined Uromonitor test detecting both TERT and FGFR3 mutations achieved accuracy of 0.902 and sensitivity of 1.00 when combined with cystoscopy. FGFR3 sensitivity is higher for high-grade tumors, while TERT shows grade-independent predictive value.
Multiple commercially available urine-based molecular tests showed clinically competitive performance for recurrence detection. The Cxbladder mRNA test targeting five genes including IGFBP5, HOXA13, MDK, CDK1, and CXCR2 achieved sensitivity of 0.93, making it valuable for ruling out recurrence in BCG-treated patients. The EpiCheck methylation test achieved accuracy of 0.883 and sensitivity of 0.917 when low-grade tumors were excluded, with performance improving substantially when restricted to high-grade NMIBC.
DNA methylation markers achieved high sensitivity for recurrence detection across multiple independent studies. The Xpert Monitor test detected five mRNA markers including ABL1, ANXA10, UPK1B, CRH, and IGF2 with accuracy of 0.79 and sensitivity of 0.74. ZNF154 methylation analysis achieved the highest sensitivity of 0.94 among individual DNA methylation markers. Three SOX1, IRAK3, and L1-MET methylation markers outperformed cytology and cystoscopy for early recurrence detection with sensitivity of 0.86.
Only two genomics studies applied AI algorithms, both showing promise but requiring external validation. Frantzi and colleagues used SVM to analyze 106 peptide biomarkers including collagen fragments and apolipoprotein sequences, achieving sensitivity of 0.88 and specificity of 0.51. Bartsch and colleagues applied genetic programming to identify rule-based ensemble classifiers from whole-genome profiling, achieving sensitivity of 0.71 and specificity of 0.69 with a three-gene rule. Both studies were limited by absence of external dataset validation.
Combined marker models consistently outperformed single-modality approaches across all study types. Nineteen studies evaluated combinations of clinical, pathological, radiomics, and genomic markers. The highest performing combination achieved accuracy of 0.975 and sensitivity of 0.966 using an MLP-based neural network combining eight clinicopathological markers with CD34 genomic expression in 308 patients treated with BCG immunotherapy.
Metaclassifier ensembles combining SVM, KNN, random forest, AdaBoost, and gradient-boosted trees showed robust performance across 1, 3, and 5-year prediction horizons. Hasnain and colleagues evaluated three models in cohorts of 2,695 to 3,071 patients. A single tumor stage marker achieved sensitivity of 0.826 for 1-year prediction but degraded at longer timeframes, while the 52-marker metaclassifier maintained balanced sensitivity around 0.7 and specificity around 0.7 across all three prediction windows.
DeepSurv, a deep neural network for survival prediction, achieved C-index of 0.651 to 0.660 for long-term recurrence prediction in 3,892 patients. Jobczyk and colleagues applied this Cox proportional hazards-inspired architecture to eight clinicopathological markers including EORTC and CUETO scores. External validation confirmed that DeepSurv could predict recurrence across different treatment arms including BCG and mitomycin C, providing a promising deep learning framework for clinical implementation.
Combined urinary genomic plus clinicopathological models achieved AUC of 0.84 for current recurrence detection. Gogalic and colleagues used LASSO logistic regression to combine five urinary markers including ECadh, IL8, MMP9, EN2, and VEGF with tumor stage, recurrence history, and BCG treatment count. The combined model with creatinine adjustment outperformed either urinary markers alone (AUC 0.75) or clinical markers alone (AUC 0.72), supporting multimodal data integration as the most accurate recurrence prediction strategy.
The majority of studies in this review used small patient cohorts and relied on internal validation only, limiting generalizability. Many included studies enrolled fewer than 200 patients, and only a minority performed external validation on independent cohorts. The heterogeneity of patient populations, imaging protocols, and outcome definitions across studies makes direct performance comparisons difficult and reduces confidence in reported accuracy metrics.
Standard clinical risk scoring systems including EORTC and CUETO showed modest C-index values of 0.55 to 0.61 when validated across three independent international cohorts. This performance floor, confirmed in a 1,892-patient validation study, demonstrates that conventional tools alone are insufficient and justifies investment in more sophisticated AI-based prediction systems. The poor calibration of existing tools across demographic groups reinforces the need for personalized approaches.
AI-based approaches in this review were predominantly applied to genomic and histopathological markers rather than radiomics or clinical data. Only two studies applied machine learning to genomic markers, and deep learning for whole-slide image analysis was evaluated in a small number of studies. This imbalance reflects the relative difficulty of acquiring standardized imaging and pathology data at scale compared to clinical variables, but limits direct comparison across biomarker categories.
Interpretability and clinical trust remain barriers to AI adoption in NMIBC surveillance. Most machine learning models in this review did not incorporate explainability tools that identify which features drove individual predictions. Clinicians require transparent reasoning to adopt algorithmic recommendations into treatment decisions, and regulatory bodies require interpretability for diagnostic AI approval. Attention visualization and feature attribution methods are underutilized in the NMIBC prediction literature.
AI-based prediction systems represent the most promising path toward personalized NMIBC management and reduced cystoscopy burden. The review identifies fusion models combining radiomics, histopathological, genomic, and clinical markers as the highest-performing category. Multimodal deep learning architectures that jointly analyze imaging, molecular, and clinical data could provide recurrence predictions surpassing any single biomarker class while enabling individualized treatment decisions.
Urine-based genomic tests with high negative predictive values offer the most immediate path to reducing invasive surveillance frequency. Tests including Cxbladder, EpiCheck, and Uromonitor demonstrated high NPV values, indicating that a negative test reliably rules out recurrence. Implementing these as first-line screening tools before deciding to perform cystoscopy could substantially reduce patient burden and healthcare costs without compromising detection rates.
Prospective studies with long follow-up are required to validate the most promising combined marker models. Most studies in this review were retrospective and did not include follow-up periods long enough to capture late recurrences occurring beyond 5 years. Prospective registries with standardized data collection across radiomics, pathology, and genomic domains would provide the high-quality training data needed for generalizable AI models.
Standardization of imaging protocols, reporting systems, and AI model evaluation frameworks is a prerequisite for clinical translation. Variability in CT and MRI acquisition, staining protocols, and outcome definitions across centers prevents aggregation of datasets needed to train robust models. Adoption of VI-RADS for imaging reporting and harmonized genomic testing platforms would accelerate multicenter AI model development and enable clinical deployment of NMIBC recurrence prediction tools that reduce the global cystoscopy burden.