The grading problem The 2021 WHO Classification introduced a three-tier grading system for invasive nonmucinous adenocarcinoma (INMA): grade 1 (well differentiated, lepidic predominant), grade 2 (moderately differentiated, acinar/papillary predominant), and grade 3 (poorly differentiated, greater than or equal to 20% high-grade pattern including solid, micropapillary, cribriform, or complex glandular types). Grade 3 INMA carries poor prognosis, higher lymph node metastasis risk, and requires lobectomy rather than sublobar resection.
Why preoperative identification matters Surgical strategy for lung adenocarcinoma depends on differentiation grade: poorly differentiated grade 3 tumors warrant lobectomy with systematic lymph node dissection and consideration of adjuvant chemotherapy, while grade 1/2 tumors may qualify for segmentectomy or wedge resection. Without preoperative grading, surgeons cannot optimally plan the extent of resection, and postoperative upstaging forces treatment adjustments after unnecessary limited resections.
Study approach Researchers from three hospitals in Jilin Province, China enrolled 451 patients with INMA confirmed by surgery. CT radiomic features from the intratumoral region, a combined intratumoral plus 3mm peritumoral region, and a combined intratumoral plus 5mm peritumoral region were extracted from non-enhanced CT scans. Clinical models and radiomic models were compared, and three radiologists evaluated cases with and without AI model assistance.
Key findings The combined 3mm radiomic model achieved the best performance with AUC 0.907 in the internal test cohort and AUC 0.772 in the external test cohort. All three radiologists - junior and senior - significantly improved their diagnostic accuracy, sensitivity, specificity, and confidence when using the combined 3mm radiomic model as decision support, with inter-reader agreement improving from moderate (kappa 0.431) to excellent (kappa 0.809).
The new IASLC grading system The International Association for the Study of Lung Cancer (IASLC) proposed a new grading scheme in 2020, subsequently adopted in the 2021 WHO classification. Grading is based on the predominant histologic pattern plus any high-grade component above a 20% threshold. High-grade patterns (micropapillary, solid, cribriform, complex glandular) in grade 3 INMA correlate with significantly worse disease-free survival and overall survival compared to grades 1 and 2.
Clinical consequences of grade 3 INMA Poorly differentiated grade 3 INMA is associated with lymph node metastasis (24.1% in grade 3 vs. 1.7% in grade 1/2 in this study's training cohort), vascular invasion (23.3% vs. 1.7%), bronchial infiltrates, and pleural involvement. These pathological features directly drive surgical decision-making: grade 3 patients require more extensive resection and should be considered for adjuvant chemotherapy based on emerging evidence.
CT imaging limitations for grading Routine CT assessment of nodule attenuation, consolidation-to-tumor ratio (CTR), and nodule features can suggest invasive components but cannot reliably distinguish grade 3 from grade 1/2 INMA before surgery. Visual assessment is subjective, experience-dependent, and shows only moderate inter-observer agreement. Radiomics offers a systematic quantitative approach to extract sub-visual texture patterns correlating with histologic grade.
Peritumoral microenvironment rationale The tumor microenvironment (TME) in the peritumoral region immediately surrounding the tumor contains biologically relevant information: tumor-infiltrating lymphocytes, macrophages, and stromal remodeling patterns that reflect aggressiveness. Prior work showed that 5mm peritumoral radiomic features distinguish benign from malignant nodules, and 3mm/5mm peritumoral models predict prognosis and EGFR mutations in NSCLC. The current study extends this concept to histologic grading.
Three-hospital cohort design A total of 451 patients (177 male, 274 female; average age 59.2 years) from the First Hospital of Jilin University, Liaoyuan Central Hospital, and Meihekou Central Hospital were enrolled retrospectively from January 2019 to July 2022. After exclusions, 413 nodules were analyzed: 289 in the training cohort (173 grade 1/2, 116 grade 3), 124 in the internal test cohort (89 grade 1/2, 35 grade 3), and 38 in the external test cohort (26 grade 1/2, 12 grade 3). All pathological diagnoses used the IASLC grading scheme reviewed by a senior pathologist with 15 years of experience.
CT acquisition and clinical feature extraction Non-contrast CT scans were acquired on Philips iCT, Siemens Cardiac 64, and GE Revolution 64 scanners at 110-120 kVp with 1.0-1.5mm reconstruction thickness. Two blinded senior chest radiologists assessed nodule attenuation (solid vs. subsolid), tumor size, consolidation size, CTR (consolidation-to-tumor ratio), and morphologic features (lobulation, spiculation, vacuolation, pleural indentation, vascular shadow). Stepwise multivariable logistic regression identified nodule attenuation, consolidation size, and CTR as independent predictors for the clinical model.
Peritumoral ROI segmentation Intratumoral ROIs were manually delineated layer-by-layer on standard lung window using RIASEG software by two blinded radiologists. The intratumoral mask was then expanded by 3mm and 5mm to create the combined peritumoral ROIs. Normal structures within the expanded regions (chest wall, ribs, blood vessels, bronchi) were manually erased. ICC above 0.75 was required for inter-observer reproducibility before features entered modeling. Combined regions are labeled 'combined 3mm' (intratumor plus 3mm peritumoral ring) and 'combined 5mm' (intratumor plus 5mm peritumoral ring).
Radiomic feature extraction and selection Using RIAS software with isotropic resampling (1x1x1mm voxels), wavelet transform, and Laplacian of Gaussian filtering, 999, 1400, and 1020 features were extracted from the intratumor, combined 3mm, and combined 5mm regions respectively. Two-stage selection retained only features with ICC above 0.75, then applied LASSO regression with 10-fold cross-validation, reducing to 8 features (intratumor), 21 features (combined 3mm), and 16 features (combined 5mm). Logistic regression models were built with 5-fold cross-validation in the training cohort.
Clinical feature differences between grades Grade 3 INMA presented predominantly as solid nodules (87.9% in training, 74.3% in internal test vs. 20-22% for grade 1/2), with significantly larger consolidation size (1.90cm vs. 0.70cm), higher CTR (0.90 vs. 0.33), and higher lobulation rates (89.7% vs. 68.8%) in training (all p less than 0.001). Pathologically, grade 3 showed markedly higher lymph node metastasis (24.1% vs. 1.7%), vascular infiltration (23.3% vs. 1.7%), and bronchial infiltration (33.6% vs. 16.8%).
Model AUC comparison in internal test In the internal test cohort (n=124), AUC values were: clinical model 0.875, intratumor radiomic model 0.882, combined 3mm radiomic model 0.907, and combined 5mm radiomic model 0.858. The combined 3mm model achieved the highest accuracy (0.855) and specificity (0.876). DeLong testing showed the combined 3mm model was statistically significantly better than the combined 5mm model (p=0.005), but differences vs. the clinical model (p=0.330) and intratumor model (p=0.166) were not statistically significant.
External test cohort performance In the external test cohort (n=38), all models showed reduced but similar performance: clinical model AUC 0.760, intratumor radiomic model 0.760, combined 3mm model 0.772, and combined 5mm model 0.766. The combined 3mm model maintained the highest specificity (0.961) but had lower sensitivity (0.500) than the clinical model (0.833), likely reflecting the small external cohort size and scanner variability across institutions.
Decision curve analysis DCA confirmed net clinical benefit from the combined 3mm radiomic model at threshold probabilities above 10% in both internal and external test cohorts, with the combined 3mm model's curve consistently highest across the probability range - outperforming both the treat-all and treat-none strategies and the other models across nearly the entire threshold range from 0.1 to 1.0.
Radiologist performance without AI Three radiologists (junior radiologist 1, junior radiologist 2, and senior radiologist 3) independently evaluated 124 internal test cases. Without AI assistance, AUC values ranged from 0.666 to 0.765, with accuracy between 0.669 and 0.750. Senior radiologist 3 significantly outperformed junior radiologist 1 (p=0.045) but not junior radiologist 2 (p=0.204), suggesting that even experienced radiologists struggle to reliably grade INMA from CT alone.
AI-assisted performance improvements With the combined 3mm radiomic model's output available, all three radiologists improved significantly (all p less than or equal to 0.005). AUC values rose to 0.821 for junior radiologist 1, 0.827 for junior radiologist 2, and 0.850 for senior radiologist 3. After AI assistance, there was no longer a significant difference between junior and senior radiologists, indicating that AI equalizes diagnostic performance across experience levels.
Inter-reader agreement transformation Without AI assistance, inter-reader consistency was only moderate (Fleiss' kappa 0.431). With AI assistance, consistency improved to excellent (kappa 0.809). This improvement reflects not only better accuracy but also greater diagnostic confidence: average confidence scores on the 5-point Likert scale improved for all three radiologists after AI consultation.
Surgical planning implications The study authors propose that patients predicted as high-risk (grade 3) by the combined 3mm model should be directed toward lobectomy with systematic lymph node dissection and postoperative adjuvant chemotherapy consideration. Patients with negative predictions for poorly differentiated INMA may be safely offered sublobar resection (segmentectomy or wedge resection), preserving lung parenchyma. This AI-guided triage aligns surgical aggressiveness with tumor biology before the operating room.
Retrospective design with selection bias The study is retrospective with data from three hospitals, introducing potential selection bias in case inclusion. Eligibility criteria (nodule under 3cm, complete clinical and pathological data, no prior treatment) may not fully reflect the real-world diversity of INMA presentations. A prospective study design with predefined enrollment and outcome assessment would provide stronger evidence for the model's clinical utility.
Manual segmentation limitations All tumor ROIs and peritumoral expansions were manually delineated by radiologists, which is time-consuming and inherently subject to inter-observer variability even with established ICC thresholds. While reproducibility was assessed for 50 lesions, the manual workflow is not scalable for routine clinical use. Development of automated or semi-automated peritumoral segmentation tools is needed for practical deployment.
Limited external cohort and scanner heterogeneity The external test cohort comprised only 38 patients from two hospitals with different CT scanners and reconstruction thicknesses. The resulting lower sensitivity (0.500) for the combined 3mm model in external testing likely reflects insufficient sample size and imaging protocol variability rather than true model failure. Larger multicenter external validation studies with standardized acquisition protocols are required.
Future directions The authors plan prospective multicenter studies with standardized data acquisition, preprocessing pipelines, and quality control systems to minimize cross-site variability. Future work will also explore deep learning-based approaches to improve classification accuracy and extend the grading model to distinguish grade 1 from grade 2 INMA - a distinction not addressed in the current study. Integration of radiomic features with molecular and genetic profiling may further reveal the relationship between imaging phenotype and tumor biology.