Clinical Problem: Pulmonary nodules detected on CT scans are common but predicting malignancy is challenging. Misclassification leads to either unnecessary invasive procedures for benign nodules or delayed treatment for malignant ones.
Study Goal: This retrospective single-center study directly compared the Lung Cancer Prediction - Convolutional Neural Network (LCP-CNN) deep learning model against the Brock statistical model and Lung-RADS classification system for malignancy risk stratification.
Population: 297 patients with 422 pulmonary nodules, of which 105 (24.9%) were malignant. The cohort was analyzed both as a whole and in clinically important subcohorts: screening, emphysema, and interstitial lung disease (ILD) patients.
Key Finding: LCP-CNN achieved superior AUC (0.92 total, 0.93 screening) compared to the Brock model (0.88), demonstrating the advantage of deep learning for nodule malignancy prediction.
Retrospective Design: The study analyzed consecutive patients with pulmonary nodules evaluated at a single center, ensuring real-world representativeness of the clinical population.
Subcohort Analysis: Four subgroups were defined - total, screening, emphysema, and ILD cohorts - to assess whether model performance varied by the underlying lung condition, which can confound nodule appearance on CT.
Malignancy Distribution: With 105 malignant nodules among 422 total (24.9%), this cohort has a higher malignancy prevalence than general screening populations, reflecting the clinical evaluation setting.
Nodule Characteristics: Both solid and subsolid nodules were included across a range of sizes, providing a realistic spectrum of the diagnostic challenge faced by radiologists and AI systems.
LCP-CNN: A commercially available convolutional neural network trained on large CT datasets to predict the probability of malignancy for pulmonary nodules based on imaging features without requiring manual feature extraction.
Brock Model: A validated multiparametric statistical model that combines clinical variables (age, sex, smoking, family history) with CT imaging features (nodule size, type, location, spiculation) to estimate malignancy risk.
Lung-RADS: The Lung Imaging Reporting and Data System provides a standardized categorical risk classification (categories 1-4) primarily based on nodule size and morphology, widely used in clinical practice.
Performance Metrics: AUC, sensitivity, and specificity were calculated for each model. DeLong's test was used to determine whether differences in AUC between models were statistically significant.
Total Cohort: LCP-CNN achieved AUC 0.92, significantly outperforming the Brock model (AUC 0.88) in the total cohort (DeLong test, p < 0.05), with Lung-RADS performing similarly to Brock.
Screening Subcohort: In the screening-specific population, LCP-CNN's advantage was maintained (AUC 0.93 vs. Brock 0.88), which is clinically significant as screening is where nodule risk stratification has the greatest public health impact.
Emphysema Subcohort: Performance differences between models narrowed in the emphysema subcohort, suggesting that background parenchymal abnormality may challenge all three approaches, including deep learning.
ILD Subcohort: In patients with interstitial lung disease, all models showed reduced performance compared to the unselected population, reflecting the additional diagnostic complexity introduced by background ILD.
AI as Decision Support: LCP-CNN's superior discrimination suggests it could improve clinical decision-making for lung nodule workup, potentially reducing unnecessary invasive procedures for benign nodules identified as low-risk.
Screening Program Integration: Given LCP-CNN's strongest advantage in the screening subcohort, integrating deep learning into LDCT screening programs could improve the positive predictive value of screen-detected nodules.
Special Populations: The reduced performance in emphysema and ILD patients highlights that AI models trained on general populations may not fully account for background parenchymal changes that alter nodule appearance.
Clinical Workflow: Rather than replacing Lung-RADS or clinical judgment, LCP-CNN could serve as an adjunct that provides a probabilistic risk estimate alongside categorical Lung-RADS assignments for borderline cases.
Single Center Retrospective Design: This retrospective single-center study may not capture the full diversity of CT protocols, scanner types, and patient populations that an AI model will encounter in real-world deployment.
Higher Malignancy Prevalence: The study population (24.9% malignancy) has a higher cancer prevalence than typical screening cohorts (1-3%), which can inflate AUC estimates compared to real-world screening scenarios.
Dynamic Nodule Assessment: All three models were applied to single time-point CT scans; incorporating longitudinal nodule growth rate data - which is highly predictive of malignancy - was not assessed.
Multicenter Prospective Validation: External validation in prospective multicenter screening cohorts with standardized CT acquisition protocols is needed to confirm that LCP-CNN's performance advantage is generalizable.