Locally advanced breast cancer is typically treated with neoadjuvant chemotherapy (NAC) followed by surgery. While patients who achieve a pathological complete response (pCR) - meaning no residive invasive cancer is found in the surgical specimen - have significantly better long-term outcomes, the majority of patients do not achieve pCR and have residual disease after NAC.
For patients who fail to achieve pCR, accurately predicting their long-term prognosis is critical. This information determines whether to escalate adjuvant therapy (adding further treatments after surgery) for high-risk patients, or de-escalate therapy for lower-risk patients who could avoid unnecessary toxicity. However, current prognostic tools have limited accuracy for this specific patient group.
Existing prognostic approaches rely on factors such as tumor stage, residual cancer burden (RCB) score (a quantitative assessment of remaining tumor after NAC), and molecular subtype. While RCB is an established independent predictor, individual factors including standard PET metabolic parameters like SUVmax have inconsistent relationships with survival, likely due to tumor heterogeneity that single metrics cannot capture.
This study developed and validated a novel integrated model combining radiomics (mathematical feature extraction from medical images) and deep learning-derived depth features from pre-treatment PET/CT scans with clinical variables, aiming to provide significantly more accurate 5-year disease-free survival prediction for non-pCR breast cancer patients.
This retrospective study analyzed 105 female non-pCR breast cancer patients (stage II-III) who underwent pre-treatment PET/CT scanning and NAC at Guangdong Provincial People's Hospital between 2009 and 2014. All patients had confirmed residual disease on surgical pathology after completing NAC and were followed for a median of 71 months (approximately 6 years).
The patient cohort had a mean age of 47 years and included diverse molecular subtypes: 62% were hormone receptor-positive/HER2-negative (luminal), 30% were HER2-positive, and 8% were triple-negative. During follow-up, 15 disease relapses and 7 deaths occurred, while 83 patients remained disease-free - representing a 21% event rate.
Regarding RCB scores, which classify residual disease burden after NAC: 14% had RCB-I (minimal residual disease), 46% had RCB-II (moderate), and 40% had RCB-III (extensive). This distribution reflects the enrolled population consisting only of non-pCR patients, with the majority having moderate-to-heavy residual disease.
Pre-treatment 18F-FDG PET/CT scans were performed using a standardized protocol (Biograph16 scanner, 7.4 MBq/kg FDG dose, minimum 6-hour fasting, 60-minute uptake period). The dataset was randomly split 70:30 into a training cohort of 73 patients and a validation cohort of 32 patients to enable model development and independent testing.
The study extracted two types of image features from the primary tumor region on PET/CT scans. Radiomic features (3,644 total) were calculated using the PyRadiomics software package following Imaging Biomarker Standardization Initiative (IBSI) standards, capturing statistical, shape, and textural properties of the tumor. Depth features (4,096 total) were extracted using the ResNet-101 deep neural network, which captures high-level image patterns that are not easily described mathematically.
Tumor regions were delineated in three dimensions using 3D-Slicer software with a semi-automatic segmentation algorithm, independently reviewed by two nuclear medicine experts. To ensure reproducibility, features with poor inter-rater agreement (interclass correlation coefficient below 0.75) were excluded, removing 1,945 unstable features from further analysis.
A rigorous multi-step feature selection process followed. First, the Mann-Whitney U test identified 1,536 features significantly associated with prognosis. Then the Boruta method - a wrapper algorithm using Shapley values - selected robust features with stronger predictive signal than random noise controls. Finally, univariate and multivariate Cox regression analyses reduced the final feature set to 4 radiomic features and 7 depth features for model building.
Five models were constructed and compared: a clinical model using tumor stage (cT) alone, a clinical model using RCB score alone, a radiomic features model, a deep learning depth features model, and an integrated combined model incorporating all clinical, radiomic, and depth features. This systematic comparison allowed assessment of the marginal contribution of each feature type.
The integrated combined model incorporating RCB, clinical tumor stage (cT), radiomic features, and deep learning depth features achieved the highest predictive accuracy across all comparisons. In the training cohort, it reached AUC values of 0.903 for 3-year survival and 0.943 for 5-year survival. In the independent validation cohort, it achieved AUC values of 0.889 for 3-year and 0.938 for 5-year survival.
The deep learning depth features model alone achieved 5-year AUC values of 0.884 (training) and 0.875 (validation), while the radiomic model achieved 0.849 and 0.806 respectively. Both clearly outperformed either clinical model alone: RCB achieved only 0.551 and 0.507 in training and validation, while the cT stage model reached only 0.683 and 0.615 - confirming that imaging-derived features provide substantially more prognostic information than routine clinical staging.
Survival stratification analysis demonstrated that the combined model successfully divided non-pCR patients into statistically distinct high-risk and low-risk groups in both the training and validation cohorts. Patients classified as high-risk by the model had significantly worse disease-free survival curves compared to low-risk patients, with the combined model showing the sharpest separation of survival outcomes among all five models.
Decision curve analysis (DCA) confirmed the clinical utility of the combined model by demonstrating higher net benefit across a broad range of threshold probabilities in both cohorts. This means the model would result in better clinical decisions (fewer missed high-risk patients, fewer over-treated low-risk patients) compared to using clinical factors alone.
A key finding of this study is that simple metabolic parameters such as SUVmax, the most commonly used PET quantitative metric, did not improve prognostic prediction when added as a standalone clinical factor. This is consistent with the known limitations of single-point uptake metrics, which capture only the peak metabolic intensity in a tumor while missing spatial heterogeneity, texture, and structural features that may be more prognostically relevant.
Tumor heterogeneity - the variation in biological and metabolic characteristics across different regions of a tumor - is a major driver of treatment resistance and recurrence. Radiomic texture features capture this heterogeneity quantitatively, while deep learning features capture additional complex spatial patterns that are computationally identified but may not have simple biological interpretations. Together, they provide complementary information that single metrics miss.
The superiority of the combined model over RCB alone is clinically significant. While RCB is a validated, widely used prognostic tool, it is calculated from the surgical specimen after completion of NAC and reflects only the immediate post-treatment pathological state. The PET/CT imaging features are obtained before NAC begins, potentially enabling earlier identification of patients who will ultimately have poor outcomes despite treatment.
The study's limitations include its retrospective, single-center design conducted at a single Chinese institution. The relatively small cohort (105 patients) and the low event rate (21%) may limit statistical power, and external validation at other centers with different patient populations and imaging protocols is needed before clinical deployment.
The primary clinical application of this model would be to identify high-risk non-pCR patients who are most likely to experience recurrence or death within 5 years despite completing standard NAC and surgery. These patients could be candidates for treatment escalation strategies, such as capecitabine or olaparib maintenance therapy, which are already guideline-recommended for certain high-risk subgroups.
Conversely, non-pCR patients classified as low-risk by the model - those with favorable imaging features despite having residual disease - might be spared from intensive adjuvant therapies whose side effects include cardiac toxicity, secondary cancers, and impaired quality of life. This enables a more individualized approach to the difficult post-NAC management question.
The model's use of pre-treatment PET/CT images is an important practical advantage. Since PET/CT is already recommended by NCCN guidelines for locally advanced breast cancer staging, these scans are routinely performed before NAC begins. The prognostic model could therefore be applied without any additional imaging burden to the patient, generating survival predictions as a byproduct of standard staging scans.
The predictive model was presented as a nomogram - a graphical calculation tool that translates patient-specific factor scores into estimated 3-year, 5-year, and 7-year survival probabilities. Nomograms are practical clinical tools that can be used at the bedside without specialized software, making the model's predictions accessible for treatment planning discussions.
This study demonstrates that an integrated model combining radiomic features, deep learning depth features, and clinical variables from pre-treatment PET/CT scans can accurately predict 5-year disease-free survival in breast cancer patients who fail to achieve pCR after neoadjuvant chemotherapy. The combined model achieved excellent AUC values of 0.943 and 0.938 in training and validation cohorts respectively.
The clear superiority of the integrated approach over any individual prognostic measure supports the emerging paradigm in precision oncology: combining multiple data streams (imaging features, clinical features, molecular profiles) into integrated models provides substantially better prediction than optimizing any single measure. This aligns with broader trends in AI-driven oncology decision support.
Future work should focus on external multicenter validation using prospectively collected data, which would test whether the model generalizes across different institutions, patient populations, and imaging equipment. Larger cohorts would also enable subgroup analyses to determine whether model performance varies by molecular subtype, which could guide subtype-specific applications.
Integrating these PET/CT-based imaging features with other emerging prognostic tools such as circulating tumor DNA, genomic signatures, and post-NAC pathological findings may further enhance prediction accuracy. The development of such comprehensive, multi-modal prognostic systems represents a promising frontier toward fully individualized breast cancer management after neoadjuvant chemotherapy.