Diagnostic Accuracy of CT for Prediction of Bladder Cancer Treatment Response with and without Computerized Decision Support

Acad Radiol 2019 AI 8 Explanations View Original
Original Paper (PDF)

Unable to display PDF. Download it here or view on PMC.

Plain-English Explanations
Pages 1-2
The Challenge of Assessing Chemotherapy Response in Bladder Cancer

Muscle-invasive bladder cancer is typically treated with neoadjuvant chemotherapy followed by radical cystectomy, a combination that improves survival and reduces metastatic spread compared to surgery alone.

However, neoadjuvant chemotherapy carries significant toxicities including neutropenic fever, sepsis, mucositis, nausea, vomiting, and hair loss. Patients who do not respond to chemotherapy bear these risks without meaningful benefit.

Currently no reliable non-invasive method exists to determine whether a patient has achieved a complete pathologic response to neoadjuvant chemotherapy. Without this information, treatment cannot be personalized and bladder-sparing alternatives (such as trimodal therapy) cannot be reliably offered.

Identifying complete responders is critically important: patients who achieve pathologic T0 disease (no residual cancer found at cystectomy) may be candidates for bladder-sparing approaches, potentially avoiding the morbidity of radical surgery.

TL;DR: No reliable non-invasive method currently exists to determine if bladder cancer has completely responded to chemotherapy, preventing personalized treatment decisions and bladder-sparing therapy selection.
Pages 3-4
Patient Cohort and CT Imaging Protocol

Pre- and post-chemotherapy CT scans from 123 patients with muscle-invasive bladder cancer were collected retrospectively, covering 157 individual cancer foci. All patients subsequently underwent radical cystectomy, providing surgical pathology as the definitive reference standard.

Chemotherapy regimens included MVAC (methotrexate, vinblastine, doxorubicin, cisplatin) and alternative regimens including carboplatin, paclitaxel, gemcitabine, and etoposide. Pre-treatment CT scans were acquired no more than 3 months before the first chemotherapy cycle, and post-treatment scans were acquired within 1 month of completing three cycles of chemotherapy.

An expert radiologist with 32 years of CT reading experience annotated all cancer foci using a bounding box on both pre- and post-treatment scans, and independently rated lesion subtlety and treatment assessment complexity on 5-point scales to enable subgroup analysis.

Pathologic T0 disease (complete response) was confirmed in 40 of 157 cancer foci (25%), with the remainder showing residual tumor at pathology stages T1 through T4.

TL;DR: The study used paired pre and post-chemotherapy CT scans from 123 patients, with surgical pathology confirming complete response in 25% of 157 cancer foci.
Pages 4-5
CDSS-T System: Deep Learning Plus Radiomics

The computerized decision support system for treatment response assessment (CDSS-T) combines two complementary analytical approaches: a deep learning convolutional neural network (DL-CNN) and a radiomic feature-based assessment model.

The DL-CNN used 32x16 pixel regions of interest extracted from within segmented tumors on both pre- and post-treatment scans. These paired ROIs were combined into 32x32 hybrid images (pre and post side by side) and used to train a classifier distinguishing complete from incomplete responders, using leave-one-case-out cross-validation to prevent overfitting.

The radiomic model extracted 91 texture and morphology features from segmented tumor volumes, calculated the percent change in each feature between pre- and post-treatment, and used a two-loop cross-validation framework to select approximately four discriminative features and train a classifier.

A final combined score was generated by taking the maximum of the DL-CNN and radiomic scores, then scaled to a 1 to 10 integer scale for display to physicians. A score distribution plot for complete and non-complete responders was shown alongside each individual score to communicate uncertainty context.

TL;DR: CDSS-T uses a hybrid of deep learning (comparing pre/post tumor images) and radiomic feature analysis, generating a 1 to 10 score shown to physicians as decision support.
Pages 5-6
Observer Performance Study Design

Twelve physician observers participated: five attending abdominal radiologists, four radiology residents, two attending oncologists, and one attending urologist. All evaluated the same 157 cancer foci independently, blinded to pathology results and each other's assessments.

Observers first estimated the likelihood of complete response (0% to 100%) based on viewing paired pre/post-treatment CT scans without access to CDSS-T scores. They then viewed the CDSS-T likelihood score and score distribution, and were given the opportunity to revise their estimates.

Multi-reader, multi-case (MRMC) receiver operating characteristic methodology was used for statistical analysis, enabling estimation of the average AUC across all observers and formal testing of whether CDSS-T produced a statistically significant improvement.

Cases were categorized as easy or difficult based on inter-reader standard deviation: treatment pairs with standard deviation 25% or less across observers were classified as easy, and those exceeding 25% were classified as difficult. This enabled separate performance analysis for diagnostically challenging cases.

TL;DR: Twelve physicians from three specialties independently assessed cancer response before and after viewing CDSS-T scores, with statistical analysis using MRMC ROC methodology.
Pages 7-8
Performance Results: Physicians vs. CDSS-T

CDSS-T alone achieved an AUC of 0.80 for identifying pathologic T0 disease, outperforming all individual physician observers who ranged from 0.66 to 0.78 without CDSS-T assistance.

The mean AUC for all 12 physicians without CDSS-T was 0.74. With CDSS-T access, mean physician AUC increased to 0.77, a statistically significant improvement (p = 0.01). Eleven of 12 individual physicians showed improvement.

Inter-observer variability also decreased significantly with CDSS-T: the mean standard deviation of likelihood estimates across observers fell from 20.4% to 17.9% (p less than 0.001), reflecting greater diagnostic consensus when AI decision support was available.

Non-radiology physicians (oncologists and urologist) showed the largest absolute gains: their mean AUC rose from 0.72 without CDSS-T to 0.76 with CDSS-T, matching the performance of radiologists using CDSS-T and achieving statistical significance (p = 0.03).

TL;DR: CDSS-T alone achieved AUC 0.80, and physician AUC significantly improved from 0.74 to 0.77 when using CDSS-T, with inter-observer variability also significantly reduced.
Pages 8-9
Easy vs. Difficult Cases and AI-Human Interaction

Subgroup analysis confirmed that CDSS-T improved performance in both easy and difficult diagnostic scenarios. For easy cases (low inter-observer variability), physician AUC improved from 0.81 to 0.84. For difficult cases (high inter-observer variability), AUC improved from 0.59 to 0.62.

The greatest reduction in inter-observer variability occurred in the difficult case subset (standard deviation decreased from 29.1% to 24.7%, p less than 0.001), indicating that physicians weighted the CDSS-T opinion more heavily when they had lower confidence in their own interpretation.

Analysis of agreement patterns showed that on average, for 8 of 40 completely responding cancers, physicians were initially wrong while CDSS-T was correct; after seeing the score, observers typically revised their estimate in the correct direction. Conversely, when CDSS-T was wrong, observers were not significantly swayed -- they maintained their original correct assessment in an average of 25 out of those cases.

This behavioral pattern suggests that physicians used CDSS-T appropriately as a second opinion, updating their assessment when the AI added new information but resisting incorrect AI guidance when their own interpretation was confident and correct.

TL;DR: CDSS-T improved physician accuracy in both easy and difficult cases, with physicians demonstrating appropriate AI deference -- following correct AI guidance while largely resisting incorrect AI suggestions.
Pages 8-10
Clinical Significance and Limitations

This study was the first to demonstrate that an AI decision support system can improve physician accuracy in assessing bladder cancer treatment response on CT, a clinically impactful but previously understudied application of computer-aided diagnosis.

The improvement was particularly meaningful for non-radiology physicians, suggesting that CDSS-T could help oncologists and urologists -- who are primarily responsible for treatment decisions -- make more accurate assessments without depending entirely on radiology interpretation.

Lesion size on pre- or post-treatment scans was not a reliable indicator of complete response, as size distributions of complete and incomplete responders substantially overlapped. This confirms that visual size-based RECIST assessment alone is insufficient, and more sophisticated image analysis is needed.

Key limitations include the use of cross-validation rather than an independent test set (due to limited data), the absence of prior CAD experience among observers which may have suppressed the magnitude of benefit, and the need to integrate additional biomarkers (genomics, proteomics, tumor resection findings) to further improve CDSS-T performance.

TL;DR: This first AI-assisted CT treatment response study showed meaningful accuracy gains especially for non-radiology physicians, though larger independent validation studies are needed before clinical deployment.
Pages 9-10
Future Directions and Bladder-Sparing Potential

Accurate identification of complete responders to neoadjuvant chemotherapy could transform treatment decision-making for muscle-invasive bladder cancer. Patients who achieve pathologic T0 disease may be candidates for bladder preservation through trimodal therapy, avoiding the substantial morbidity of radical cystectomy.

The authors plan to expand CDSS-T by integrating clinical biomarkers from transurethral resection findings, bimanual exam under anesthesia, and molecular biomarkers including genomics and proteomics to improve discrimination beyond what CT imaging alone can achieve.

Future studies should evaluate CDSS-T on independent test sets with larger cohorts, include observer training sessions to familiarize physicians with AI-assisted reading, and assess whether the diagnostic gains translate into improved patient management and outcomes in prospective clinical settings.

This type of AI decision support could be particularly valuable in community settings where radiologists may have less experience reading treatment response CT scans, enabling more consistent quality of oncologic imaging interpretation across institutions.

TL;DR: CDSS-T could enable non-invasive identification of complete bladder cancer responders who may avoid radical surgery, with future versions planned to integrate genomic and clinical biomarkers for greater accuracy.
Citation: Open Access, 2019. Available at: PMC6510656.