CT-based deep learning radiomics nomogram for the prediction of pathological grade in bladder cancer: a multicenter study.

Cancer Imaging 2023 AI 7 Explanations View Original
Original Paper (PDF)

Unable to display PDF. Download it here or view on PMC.

Plain-English Explanations
Pages 1-2
Why Preoperative Grade Prediction Matters in Bladder Cancer

Tumor grade shapes treatment decisions. Pathological grade is among the most important prognostic factors in bladder cancer. Low-grade tumors carry a 4 percent progression rate and a 43 percent recurrence rate, while high-grade tumors carry a 19 percent progression rate and a 58 percent recurrence rate. Grade determines whether a patient needs only transurethral resection or must be considered for partial or radical cystectomy.

Current grading methods have limitations. Cystoscopic biopsy, the standard grading approach, is invasive, carries risks of bladder perforation, and is prone to sampling error due to tumor heterogeneity. Biopsy specimens sometimes underestimate tumor grade, leading to undertreated high-risk disease. A noninvasive preoperative method to accurately assess grade would greatly improve patient management.

CT imaging as a noninvasive alternative. CT urography (CTU) is routinely performed for preoperative evaluation of bladder cancer. However, visual assessment of tumor heterogeneity on CT is unreliable. Radiomics, which extracts large numbers of quantitative features from CT images, can objectively characterize tumor texture and heterogeneity in ways that the naked eye cannot.

Deep learning adds value beyond radiomics. While handcrafted radiomics features are selected based on predefined mathematical formulas, deep learning automatically learns image representations that may capture different and complementary information. Combining both approaches in a multicenter study with external validation represents a significant advance over prior single-center, single-modality studies.

TL;DR: Preoperative grade determination in bladder cancer is clinically critical but limited by invasive biopsy sampling errors, motivating the development of a noninvasive CT-based deep learning radiomics nomogram.
Pages 2-3
Multicenter Cohort and CT Image Acquisition

Patient selection across three centers. A total of 688 bladder cancer patients who underwent surgical resection at three hospitals in China were enrolled. The training cohort comprised 469 patients from the Affiliated Hospital of Qingdao University, while 219 patients from Shandong Provincial Hospital and Puyang Oilfield General Hospital formed the external test cohort. Patients who received preoperative treatment or had multiple lesions were excluded.

Three-phase CT urography protocol. All patients underwent CTU with contrast enhancement. Images were acquired at the corticomedullary phase (25 seconds), nephrographic phase (75 seconds), and excretory phase (300 seconds) after bolus injection. Using all three phases rather than a single phase captured a comprehensive picture of tumor enhancement dynamics across the imaging time course.

Region of interest segmentation. Two radiologists unaware of pathological findings independently delineated the 3D tumor region of interest on each phase using ITK-SNAP software. Intra- and inter-rater agreement was assessed using intraclass correlation coefficients (ICCs) on a subset of 94 lesions re-segmented by both readers and by one reader after a 3-week interval.

ComBat harmonization for multicenter compatibility. Because CT scanner models, protocols, and parameter settings differed across centers, a ComBat compensation methodology was applied to remove systematic differences in radiomic features attributable to scanner variation. All features were subsequently normalized via Z scoring before model training.

TL;DR: Six hundred eighty-eight patients from three hospitals underwent three-phase CT urography, with ComBat harmonization applied to minimize scanner-related feature differences across centers.
Pages 3-4
Feature Extraction and Machine Learning Model Construction

Handcrafted radiomics feature extraction. PyRadiomics was used to extract 3,948 handcrafted radiomics (HCR) features from the three-phase CT images, encompassing first-order statistics, shape features, and texture features including gray-level co-occurrence matrix, run-length matrix, size zone matrix, dependence matrix, and neighboring grey tone difference matrix features, as well as wavelet-transformed versions of all these.

Deep learning feature extraction via ResNet18. A ResNet18 network pretrained on ImageNet was used as the deep learning feature extractor. The network was fine-tuned on each CT phase separately, and the output of the penultimate layer was used to define 1,536 deep learning features across the three phases per patient.

Feature selection pipeline. Only HCR features with ICC greater than 0.80 were retained (2,636 of 3,948), ensuring reproducibility. LASSO regularized regression then reduced the combined HCR plus deep learning feature set to the 60 most informative features, which included 25 HCR features and 35 deep learning features. Deep learning features carried the largest predictive weights.

Eleven machine learning classifiers evaluated. The selected features were used to train 11 different classification algorithms including logistic regression, SVM, KNN, Random Forest, XGBoost, LightGBM, and others, each assessed by 10-fold cross-validation AUC and accuracy. This broad evaluation was designed to identify the most reliable classifier for the combined feature set.

TL;DR: From 3,948 handcrafted radiomics features and 1,536 deep learning features, LASSO selected 60 final features (35 deep learning, 25 HCR) used to train 11 different machine learning classifiers.
Pages 5-6
Clinical Model and Radiomic Signature Performance

Clinical model construction. Multivariate logistic regression identified age, size, shape, boundary, stalk, and extramural infiltration as independent predictors of bladder cancer grade. The resulting clinical model achieved AUC values of 0.752 in the training cohort and 0.745 in the external test cohort, reflecting moderate discriminative ability from visual CT and demographic features alone.

Best machine learning classifier: SVM. Among all 11 classifiers tested with combined deep learning plus HCR features, the support vector machine (SVM) achieved the best and most generalizable performance: AUC of 0.953 in the training cohort and 0.943 in the external test cohort, with accuracy of 0.902 and 0.840 respectively. Tree-based ensemble methods such as Random Forest and XGBoost overfit to the training set with training AUCs of 0.999 to 1.000 but external test AUCs dropping to 0.778 to 0.850.

Deep learning features outperform HCR alone. The best model using only deep learning features (NaiveBayes, AUC 0.882 in the external test) outperformed the best model using only HCR features (MLP, AUC 0.848 in the external test), confirming that the deep learning approach captured additional grade-relevant information beyond what hand-crafted texture features alone could provide.

Wavelet features dominated the HCR contribution. Among the 25 selected handcrafted radiomics features, 20 (80%) were wavelet-derived features, with the wavelet-based neighboring grey tone difference matrix feature carrying the largest individual weight. Wavelet features capture multi-scale texture heterogeneity that is not visible in original intensity images.

TL;DR: The SVM classifier trained on combined deep learning and handcrafted features achieved AUC values of 0.953 and 0.943 in training and external test cohorts, outperforming HCR-only and DL-only models.
Pages 6, 7, 9, 10
Nomogram Performance and Survival Stratification

DLRN construction and performance. The deep learning radiomics nomogram (DLRN) combined the SVM radiomic signature with independent clinical predictors (age, size, shape, boundary, stalk, extramural infiltration). DLRN achieved AUC values of 0.961 and 0.947 in the training and external test cohorts respectively, with accuracy of 0.919 and 0.854, outperforming both the clinical model and the optimal machine learning model alone.

Calibration and clinical utility confirmed. Calibration curves demonstrated good agreement between DLRN-predicted probabilities and observed outcomes in both cohorts. Decision curve analysis showed that DLRN provided greater net clinical benefit than either the clinical model alone or treating all patients as high-grade across the clinically relevant range of threshold probabilities.

Survival stratification by DLRN. Kaplan-Meier analysis showed that DLRN successfully stratified patients by progression-free survival (PFS) in the total cohort and the training cohort. The pathological grade model and DLRN both demonstrated significant PFS differences in these groups, supporting the clinical relevance of grade prediction for long-term patient management.

Limitation in external test PFS stratification. DLRN did not significantly stratify PFS in the external test cohort alone, likely because follow-up in that cohort was shorter with no patient reaching 50 months of follow-up. The total and training cohorts better reflected the long-term PFS distribution and showed significant stratification.

TL;DR: The DLRN achieved AUC values of 0.961 and 0.947 in training and external test cohorts, with good calibration and decision curve benefit, and stratified patient survival in the overall cohort.
Pages 9-10
Imaging Interpretation and Feature Insights

Activation map visualization. CNN activation maps showed that the deep learning model focused on regions associated with higher-grade tumors, with red activation areas indicating stronger correlations with high-grade pathology. These maps demonstrated that the network focused on clinically meaningful tumor regions across all three CT phases.

Why wavelet features are informative. Wavelet-based radiomics features decompose images at multiple frequency scales, enabling quantification of texture heterogeneity that spans both fine-grained and coarse structural patterns within the tumor. Prior studies have confirmed the value of wavelet features in tumor grade classification across multiple cancer types.

Comparison with prior studies. Prior single-center CT and MRI radiomics studies achieved AUC values of 0.860 to 0.956 for bladder cancer grading. The current multicenter DLRN achieved AUC of 0.947 in the external test cohort, matching or exceeding these benchmarks while using a much larger cohort with prospectively defined external validation.

Clinical context of grade misclassification. Biopsy-based grade assignment can underestimate true tumor grade due to specimen inadequacy and heterogeneity. If a preoperative CT-based model correctly identifies high-grade tumors classified as low-grade by biopsy, it could prompt earlier escalation to more aggressive treatment and prevent disease progression in those patients.

TL;DR: CNN activation maps confirmed the model focused on meaningful tumor regions, and deep learning features outweighed handcrafted features in the final model, reflecting genuine grade-related image information.
Pages 10-11
Study Limitations and Future Directions

Demonstrated clinical utility. The CT-based DLRN demonstrated strong performance for preoperative prediction of bladder cancer pathological grade, outperforming both conventional radiomics and clinical models. It could serve as a noninvasive decision-support tool to guide treatment planning before surgery.

Retrospective design and selection bias. As a retrospective study, this work may be affected by uncontrolled selection factors. Prospective validation in diverse patient populations will be necessary to confirm that model performance generalizes beyond the institutions and time periods included in this study.

Manual segmentation dependency. The pipeline required manual 3D tumor delineation by trained radiologists, which is time-consuming and subject to inter-observer variability. Integrating automatic segmentation algorithms would be a necessary step toward routine clinical deployment.

Future enhancements. The study included only baseline clinical and CT data and did not incorporate MRI, molecular markers, or other biomarker data. Future models integrating multi-parametric MRI, genomic data, and automated segmentation may further improve preoperative grade prediction accuracy and clinical applicability.

TL;DR: The CT-based DLRN provides strong preoperative grade prediction for bladder cancer with good external validation, though prospective studies, automatic segmentation, and multi-modal integration remain important next steps.
Citation: Open Access, 2023. Available at: PMC10507832.