Lung adenocarcinoma (LUAC) is the most common subtype of non-small cell lung cancer and can be classified into five histological patterns: lepidic, acinar, papillary, micropapillary (MP), and solid (S). These patterns matter enormously because patients whose tumors contain MP or solid components have substantially worse survival outcomes than those with other subtypes.
Even when MP or solid patterns are not the predominant tumor component - constituting as little as 5% of the tumor - their presence is still associated with poorer prognosis, higher rates of lymph node metastasis, lymphovascular invasion, and tumor recurrence. This makes preoperative identification of these patterns clinically critical.
The surgical stakes are high. For most early-stage LUAC, minimally invasive sublobar resection (segmentectomy or wedge resection) can be appropriate. However, for tumors with MP/S components, lobectomy with comprehensive lymph node dissection is preferred because limited resection carries a higher risk of leaving residual tumor cells and subsequent recurrence.
Currently, no accurate preoperative model exists to guide surgical decision-making for patients with MP/S pattern LUAC. This study aimed to fill that gap by developing and validating a machine learning model using standard clinical and CT imaging variables that are available before surgery.
This retrospective study analyzed 1,933 patients with pathologically confirmed stage IA lung adenocarcinoma who underwent surgical resection at Sun Yat-sen University Cancer Center. A total of 14.43% of patients had MP/S positive tumors (279 patients), and propensity score matching at a 1:2 ratio was used to create a balanced comparison group of 558 MP/S negative patients.
Propensity score matching is a statistical technique that creates comparable groups by matching patients on observable characteristics such as age, sex, tumor size, and comorbidities - mimicking the balance that randomization achieves in clinical trials. This reduces confounding that could otherwise bias comparison of MP/S positive versus negative patients.
The cohort was divided chronologically: patients from January 2022 to March 2023 formed the internal cohort (split 70:30 into training and validation sets), while patients from April 2023 to August 2023 formed the external validation cohort. This chronological split ensures the model is tested on patients encountered after training - a realistic assessment of real-world applicability.
CT morphological features were assessed by two specialized radiologists blinded to pathological results, with intraclass correlation coefficients of 0.85 to 0.96 for continuous measurements and Cohen's Kappa values of 0.77 to 0.91 for categorical features - indicating high inter-rater reliability and providing a reliable data foundation for modeling.
Variable selection used a two-step process. Univariate logistic regression first identified candidate predictors with p below 0.05. These were then entered into LASSO (Least Absolute Shrinkage and Selection Operator) regression, which penalizes model complexity and shrinks uninformative coefficients to zero, ultimately selecting 6 key predictors from an initial pool of 13 candidates.
The six selected predictors were: nodule type (pure ground-glass, part-solid, or solid nodule on CT), spiculation (spiky irregular margins on CT), serum carcinoembryonic antigen (CEA) level, maximum solid component diameter, median CT value, and CT value range. These represent a clinically accessible combination of imaging characteristics and a simple blood test.
Ten machine learning algorithms were evaluated: logistic regression, support vector machine (SVM), gradient boosting machine (GBM), neural network, random forest, XGBoost, K-Nearest Neighbors (KNN), AdaBoost, LightGBM, and CatBoost. Ten-fold cross-validation was used for model training, and hyperparameters were tuned using grid search combined with cross-validation.
Model comparison used multiple metrics including AUC, accuracy, sensitivity, specificity, positive and negative predictive value, and F1-score. Calibration curves assessed agreement between predicted and observed outcomes, and decision curve analysis quantified clinical net benefit at different threshold probabilities.
Solid nodule type on CT was the strongest single predictor of MP/S pattern presence in univariate analysis, with an odds ratio of 18.18 compared to pure ground-glass nodules. Part-solid nodules also showed elevated risk (OR 1.86) but the effect was smaller and did not reach statistical significance in multivariate analysis after adjusting for other factors.
Spiculation (irregular spiky tumor margins on CT) was the only CT morphological feature retained as an independent predictor in multivariate analysis, with an odds ratio of 2.08 - indicating that spiculated tumors have more than twice the risk of containing MP/S components compared to smooth-margined tumors.
Higher median CT value and narrower CT value range were independently associated with MP/S pattern presence. A higher CT value indicates greater lesion density and more solid components, which correlates with the known CT appearance of MP/S adenocarcinomas. The CT value range reflects internal heterogeneity of the nodule.
Among the ten models evaluated, the KNN (K-Nearest Neighbors) algorithm was selected as the final model based on its combination of strong discriminative performance and, crucially, its ability to maintain consistent performance between the training cohort and validation cohort - indicating good generalization rather than overfitting. Random Forest, LightGBM, and GBM all outperformed KNN in training but showed significant performance drops in validation.
The KNN model achieved an AUC of 0.871 (95% CI: 0.838-0.904) in the internal training cohort and 0.787 (95% CI: 0.717-0.857) in the internal validation cohort - indicating strong discriminative ability that was maintained between training and held-out validation data without evidence of significant overfitting.
Calibration was excellent: Hosmer-Lemeshow test p-values greater than 0.05 (0.053 for training, 0.817 for internal validation) indicated no statistically significant difference between predicted probabilities and observed outcomes across risk groups. The calibration curves closely approximated the ideal diagonal line.
External validation on the temporally separate cohort yielded an AUC of 0.790 with a Hosmer-Lemeshow p-value of 0.120, further confirming the model's generalizability. The performance remained consistent with internal validation results, providing reassurance that the model captures real clinical patterns rather than dataset-specific noise.
Decision curve analysis confirmed net clinical benefit across a range of threshold probabilities, meaning that acting on the model's predictions would benefit more patients than treating all or treating none - the standard for evaluating whether a predictive model has genuine clinical utility beyond statistical performance alone.
SHAP (SHapley Additive exPlanations) analysis was used to open the 'black box' of the KNN model and quantify each feature's contribution to individual predictions. This addresses one of the main critiques of machine learning in medicine: that models cannot explain their reasoning to clinicians.
Nodule type was the dominant predictor, with an absolute average SHAP value of 13.6%, meaning that nodule type alone can shift the predicted probability of MP/S presence by up to 13.6 percentage points. The direction of effect aligned with expectation: solid nodules strongly increased predicted probability, while pure ground-glass nodules decreased it.
Other features contributed as follows: median CT value shifted predictions by approximately 8.6%, CT value range by 6.5%, maximum solid component diameter by 3.5%, spiculation by 3.0%, and CEA by 2.8%. Notably, samples with higher CT value range (reflecting more internal heterogeneity) reduced predicted MP/S probability, while lower CT value range increased it - consistent with the more uniform, dense appearance of MP/S tumors.
SHAP provides actionable clinical transparency. Clinicians can see not just whether a prediction is high or low risk, but which specific imaging or laboratory features in that particular patient most influenced the prediction - enabling them to scrutinize borderline predictions and incorporate clinical context that the model cannot capture.
The clinical stakes of correctly identifying MP/S patterns preoperatively are substantial. Multiple large studies have shown that for patients with MP or solid components, sublobar resection (segmentectomy or wedge resection) is associated with significantly higher recurrence rates compared to lobectomy. In one study, limited resection in patients with any MP component of 5% or more resulted in substantially worse outcomes.
The pivotal JCOG0802 and CALGB140503 trials demonstrated that sublobar resection is non-inferior to lobectomy for peripheral NSCLC up to 2 cm - but these trials did not stratify by histological subtype. This study addresses that gap by enabling pathological subtype identification before surgery, allowing the trial findings to be applied to the appropriate patient subgroups.
Lymph node management also differs. For MP/S pattern patients, complete lymph node dissection significantly improved overall and recurrence-free survival compared to incomplete dissection. Patients who received inadequate lymphadenectomy had more frequent distant metastases and worse outcomes. This further underscores the importance of identifying MP/S patterns before finalizing the surgical approach.
The model uses only standard clinical variables readily available in routine preoperative workup - no specialized radiomic post-processing tools or additional imaging beyond standard CT are required. This makes it directly applicable in primary care and community settings without specialized infrastructure, broadening its potential clinical reach.
This study successfully developed and validated an interpretable KNN-based machine learning model to predict MP/S pattern presence in stage IA lung adenocarcinoma using six preoperative clinical and CT variables. The model demonstrated consistent performance across internal and external validation cohorts and provides SHAP-based explanation for individual predictions.
The practical advantage over prior approaches is the use of clinically actionable variables rather than high-dimensional radiomic features requiring specialized tools. This model can be applied within standard clinical workflows without additional software, making it a more realistic candidate for real-world adoption.
Key limitations include the retrospective, single-center design at a major Chinese cancer center, which may limit generalizability. The study population reflects patients who already underwent surgery, introducing selection bias. Multi-center prospective validation with larger samples is needed to confirm generalizability and refine performance.
Future research will focus on establishing a prospective cohort tracking survival outcomes of stage IA MP/S patients following different surgical approaches, which will allow direct validation of whether preoperative model-guided surgery selection improves patient outcomes compared to current practice.