Accurate mediastinal lymph node staging in NSCLC is critical because N0/1 patients are candidates for primary surgery while N2/3 patients typically require non-surgical or multimodal treatment, making incorrect staging directly harmful to patient outcomes. FDG-PET/CT is the guideline-recommended standard for this assessment, but its diagnostic accuracy is fundamentally limited by false-positive findings that arise from the tracer's non-specificity for malignant tissue.
The primary source of false positives is FDG uptake by non-malignant glucose-metabolizing cells such as macrophages in inflammatory conditions including anthracosis, pulmonary emphysema, and epithelioid cell reactions. Benign inflammatory lymph node enlargement can produce PET signal indistinguishable from metastatic involvement when using standard visual criteria alone. These false positives lead to unnecessary invasive confirmation procedures like EBUS-TBNA or mediastinoscopy, exposing patients to procedural risks without benefit.
The standard PET/CT criterion classifies a lymph node as positive if mediastinal uptake exceeds background activity or if short-axis diameter exceeds 10 mm, a simple rule that does not leverage the full quantitative information available from modern PET/CT acquisitions. Machine learning models that integrate multiple quantitative parameters simultaneously can potentially identify more complex patterns that distinguish true metastatic involvement from inflammatory uptake, improving specificity without sacrificing sensitivity.
A gradient boosting machine learning classifier was previously developed using 10 routinely obtainable PET/CT and clinical parameters to improve mediastinal lymph node staging in NSCLC, demonstrating significantly higher specificity (73% vs. 54%) at similarly high sensitivity compared to standard visual interpretation. This independent validation study tested that pre-trained classifier without modification on two new cohorts to assess whether performance was maintained across different institutions, patient populations, and PET/CT scanners.
The study analyzed 211 NSCLC patients across two independent cohorts: 87 patients prospectively enrolled at Charite-Universitatsmedizin Berlin (the Charite cohort) and 124 patients from the publicly available NSCLC Radiogenomics dataset within The Cancer Imaging Archive (the TCIA cohort). The two cohorts differed substantially in patient characteristics: the Charite cohort included both surgical and non-surgical patients from routine clinical staging, while the TCIA cohort contained only surgically treated patients who had been preoperatively staged as candidates for surgery. N2/3 prevalence differed significantly: 40% in the Charite cohort versus 12% in the TCIA cohort.
The previously trained machine learning classifier was applied to both cohorts without any modification, using the same probability threshold of 0.19 established during original model development to distinguish N2/3-positive from N0/1-negative classifications. The classifier uses 10 parameters: maximum standardized uptake values (SUVmax) of N1 and N2 lymph nodes (both uncorrected and background-corrected), a 4-level visual PET score for mediastinal uptake intensity, short-axis diameters of N1 and N2 lymph nodes, primary tumor diameter, and patient age. These parameters are all obtainable during routine clinical interpretation without specialized software.
Two different PET/CT scanner types were used in the Charite cohort, and the TCIA cohort was acquired at two different US institutions with yet different scanners, creating a technically heterogeneous validation environment that tests the classifier's robustness across acquisition protocols. No harmonization of imaging protocols was performed, meaning the classifier was tested under realistic conditions of cross-institutional deployment where scanner differences, reconstruction algorithms, and acquisition parameters all varied.
The reference standard was histological N status confirmed by surgery with systematic lymphadenectomy in 176 of 211 patients, with the remaining 35 Charite patients having N status confirmed via EBUS-TBNA. As comparator, the standard PET/CT criterion (mediastinal uptake exceeding background and/or lymph node short-axis exceeding 10 mm) was applied to all patients. Statistical comparison used McNemar's test, and decision curve analysis quantified net clinical benefit across a range of threshold probabilities.
In the Charite cohort, both the machine learning classifier and the standard PET/CT criterion achieved equally high sensitivity of 97.1% for detecting N2/3 disease, with no statistically significant difference in specificity (65% vs. 60%, p=0.5) or accuracy (78% vs. 75%). The similar performance in this mixed surgical and non-surgical cohort suggests that the classifier replicates the diagnostic value of standard interpretation while offering marginal specificity gains that did not reach statistical significance at this sample size.
In the TCIA cohort of exclusively surgically treated patients, the machine learning classifier achieved significantly higher specificity than the standard PET/CT criterion (90% vs. 70%, p less than 0.001), resulting in significantly higher overall accuracy (82% vs. 65%, p less than 0.001). This 20-percentage-point improvement in specificity translates directly to fewer false-positive classifications that would otherwise prompt unnecessary invasive mediastinal staging in patients already deemed candidates for primary surgery.
Sensitivity in the TCIA cohort was low for both methods (27% for the classifier, 33% for standard criterion), reflecting the inherent patient selection bias of this surgical cohort rather than classifier failure. All TCIA patients had undergone standard preoperative workup including FDG-PET/CT and been deemed N0/1 candidates for surgery before enrollment. The remaining N2/3 cases therefore represent occult metastases that had already evaded standard staging, making any method's detection in this population intrinsically limited.
Brier scores were similar and low in both cohorts (Charite: 0.11, TCIA: 0.097), indicating well-calibrated probability outputs that accurately reflect the actual likelihood of N2/3 disease. Decision curve analysis confirmed consistently higher net clinical benefit of the machine learning classifier compared to the standard PET/CT criterion in both cohorts, supporting its use across different clinical contexts despite the variation in absolute performance metrics driven by differing disease prevalence.
A key distinguishing feature of this classifier compared to high-dimensional radiomics models is its reliance on 10 clinically intuitive parameters that can be measured during routine image interpretation without specialized software or post-processing pipelines. Radiomic features derived from gray-level matrices of PET images are theoretically richer in information but are inherently susceptible to distortion from variations in scanner hardware, acquisition protocols, reconstruction algorithms, and image post-processing. The simpler parameters used in this classifier are more robust to these technical sources of variability.
The consistent performance across multiple PET/CT scanner types and acquisition protocols, including scanners at different institutions in both Germany and the United States, empirically demonstrates this robustness advantage. Despite expected differences in SUV values between scanner systems, the overall classifier predictions were rarely altered to a clinically relevant degree, suggesting that the multi-parameter model captures relative patterns within each patient's imaging data rather than absolute SUV values that vary systematically between scanners.
Comparison with other machine learning models validated on the NSCLC Radiogenomics dataset shows that the classifier's 82% accuracy is competitive with approaches using more complex deep learning architectures. A residual neural network using only CT data without PET information achieved similar 81% accuracy on the same dataset. A graph neural network combining PET and CT data achieved 83% accuracy but used higher-dimensional features requiring more complex image analysis. The current model achieves comparable performance while maintaining practical deployability.
The clinical implication of improved specificity is that some patients classified as N2/3 positive by standard PET/CT criteria but reclassified as N0/1 by the classifier could safely proceed to primary surgery without invasive mediastinal staging, reducing procedural risks and costs. This is most valuable in patients with lymph nodes that are anatomically difficult to sample by EBUS-TBNA or mediastinoscopy, in multimorbid patients with elevated procedural risk, and in patients with low pretest probability of true N2/3 disease who nonetheless have borderline PET findings.
The primary clinical use case for the classifier is in patients where standard PET/CT interpretation raises concern for N2/3 disease but clinical probability is uncertain, and the classifier provides a low-probability output that supports foregoing invasive confirmation. A probability below the 0.19 threshold indicates the classifier predicts N0/1 status, potentially allowing the multidisciplinary tumor board to proceed with surgical planning without EBUS-TBNA or mediastinoscopy, particularly when those procedures carry elevated risk for that specific patient.
The classifier is available as an openly accessible web application that accepts the 10 required parameters and outputs a probability of N2/3 disease, requiring no institutional installation or specialized computing infrastructure. The input parameters (SUVmax values, lymph node diameters, tumor diameter, age, visual PET score) are all documented in routine nuclear medicine reports, making data entry practical without additional imaging reanalysis. This accessibility positions the tool as a decision support aid that any clinician with access to the PET/CT report can use.
The radiomics quality score of 24 out of 36 (67%) achieved by this study substantially exceeds the average of 9.4/36 (26%) observed in reviews of oncological imaging studies, reflecting the methodological rigor of the independent validation approach. High RQS reflects appropriate sample size, external validation, and transparent reporting, all of which are required for a computational tool to be taken seriously as a candidate for clinical adoption. Meeting these standards distinguishes this model from the majority of published radiomics and machine learning studies that report training performance without independent validation.
Future validation should include prospective interventional studies where clinical decisions are actually influenced by classifier output, enabling measurement of real-world impact on the frequency of invasive staging procedures, staging accuracy, and ultimately patient outcomes. While the current study validates diagnostic performance in held-out cohorts, it cannot demonstrate that using the classifier in routine practice will change clinical decisions or improve outcomes without a prospective trial design. This is the standard that regulatory bodies and clinical guideline committees require for meaningful clinical adoption.
This study successfully validates a machine learning classifier for mediastinal lymph node staging in NSCLC across two independent cohorts with different disease prevalence, patient selection criteria, and PET/CT scanner types, confirming that the previously reported diagnostic performance generalizes beyond the original training environment. The classifier maintained high sensitivity comparable to standard PET/CT interpretation while achieving significantly superior specificity in the surgically staged TCIA cohort, with consistent net clinical benefit demonstrated by decision curve analysis in both validation populations.
The key methodological contribution of this validation study is demonstrating that a classifier using clinically interpretable, scanner-robust parameters can outperform standard visual interpretation in cross-institutional settings where complex radiomic features would be most susceptible to technical variability. To the authors' knowledge, this is the first machine learning model for NSCLC lymph node staging to undergo independent validation across a comparable number of independent datasets and PET/CT scanner types, setting a methodological standard for the field.
The substantial differences observed between the preselected TCIA surgical cohort and the real-world Charite clinical population underscore how profoundly patient selection affects apparent model performance, making external validation in demographically distinct populations essential before drawing conclusions about clinical utility. A model that performs well only in populations similar to its training cohort provides limited guidance for clinical adoption, while a model validated across heterogeneous populations provides much stronger evidence of practical generalizability.
Next steps include extending validation to additional independent centers and external readers, followed by a prospective interventional trial to determine whether classifier-assisted staging reduces unnecessary invasive procedures and improves treatment planning accuracy in routine clinical practice. Until such evidence is available, the classifier should function as a decision support tool that informs but does not replace physician judgment in the multidisciplinary tumor board setting.