Diagnostic Accuracy of Artificial Intelligence Models for Differentiation of Squamous Cell Carcinoma and Adenocarcinoma of Lung - A Systematic Review

Diagnostics (Basel) 2026 AI 8 Explanations View Original
Original Paper (PDF)

Unable to display PDF. Download it here or view on PMC.

Plain-English Explanations
Pages 1-2
Why Distinguishing Lung Cancer Subtypes Matters

Not all lung cancers are the same, and the distinction between subtypes has direct treatment implications. Non-small cell lung cancer (NSCLC) accounts for 80-85% of all lung cancer cases and is divided into two major subtypes: adenocarcinoma (ADC) and squamous cell carcinoma (SCC). While both arise in the lung, they respond very differently to targeted therapies, chemotherapy regimens, and have different patterns of spread and prognosis.

Adenocarcinoma tends to occur in the outer parts of the lung, is more common in non-smokers and women, and more frequently harbors targetable genetic mutations such as EGFR, ALK, and ROS1 that can be treated with highly effective oral targeted therapies. Squamous cell carcinoma more often arises centrally near major airways, is strongly associated with smoking, and generally has fewer targetable mutations but responds to different immunotherapy and chemotherapy combinations.

While tissue biopsy remains the gold standard for subtype classification, it is invasive, carries procedural risks (bleeding, collapsed lung), and may not always be feasible. Additionally, biopsies sample only a small portion of a tumor, potentially missing important spatial heterogeneity. These limitations create a strong clinical need for accurate, non-invasive methods to characterize lung tumor histology from CT imaging.

TL;DR: Distinguishing adenocarcinoma from squamous cell carcinoma is critical for treatment selection, but current biopsy-based diagnosis is invasive and limited, motivating AI-based non-invasive classification from CT images.
Pages 2-3
Systematic Review Design and Study Selection

This systematic review followed PRISMA standards to ensure rigorous and transparent reporting. A comprehensive literature search across three major databases - PubMed, Scopus, and Embase - retrieved 2,095 initial studies. After removing duplicates and applying strict inclusion and exclusion criteria, 11 final studies were selected for review.

The inclusion criteria required studies to: use AI models (machine learning or deep learning) to differentiate between lung SCC and ADC, include only patients with pathologically confirmed lung carcinoma, and report quantitative performance metrics such as accuracy, sensitivity, specificity, and AUC. Studies lacking full text, focused only on detection rather than subtype classification, or not employing radiomics or AI approaches were excluded.

Study quality was assessed using the Radiomics Quality Score (RQS) tool, a 16-item checklist designed specifically to evaluate the methodological rigor and clinical relevance of radiomic studies. The RQS evaluates items including data reproducibility, feature stability testing, validation strategy, and whether the study included cost-effectiveness analysis or prospective design. All 11 included studies scored below 50% on the RQS, indicating substantial room for methodological improvement across the field.

TL;DR: After screening over 2,000 studies from three databases, 11 qualifying AI studies were identified and assessed for quality using the specialized Radiomics Quality Score tool.
Pages 6-7
Radiomics and Deep Learning Approaches Used

The 11 included studies employed two broad approaches to CT-based lung cancer subtyping: classical radiomics (extracting handcrafted quantitative features from tumor regions, then applying machine learning classifiers) and deep learning (training neural networks directly on CT images to learn discriminative features automatically). Some studies combined both approaches.

Classical radiomics models extracted features describing tumor shape, texture, and intensity distribution - up to 3,078 features in one study - then used machine learning algorithms including Random Forest, Support Vector Machine, Naive Bayes, and k-nearest neighbors for classification. These models require explicit tumor segmentation but are more interpretable than deep learning approaches.

Deep learning models employed diverse architectures including standard convolutional neural networks (CNNs), VGG16 (a well-established image classification network), 3D CNNs (processing the full tumor volume), long short-term memory networks (LSTM, capable of processing sequential image slices), Capsule Networks (CapsNet, designed for better spatial relationship modeling), and combinations of these approaches. Each architecture offers different tradeoffs between accuracy, data requirements, and interpretability.

The total patient cohort across all studies exceeded 4,300 cases, from institutions in eight countries including China, India, Japan, Brazil, the United Kingdom, and South Korea. Studies used both institutional scanner data and public datasets (LUNA16, TCIA), providing geographic and demographic diversity that helps evaluate whether AI models generalize across populations.

TL;DR: Studies used both handcrafted radiomics with classical machine learning and end-to-end deep learning approaches, across more than 4,300 patients from eight countries and multiple public and private datasets.
Pages 7-9
Performance of AI Models Across Studies

Deep learning models achieved the highest overall performance, with accuracy ranging from 67% to 97% across studies. The top-performing system (LCDCS by Rawat et al.) achieved 96.9% validation accuracy with 95.9% sensitivity and 97.9% specificity for NSCLC detection using multi-level CNNs. Lima et al.'s VGG16-based system achieved 84.5% accuracy for adenocarcinoma and 89.6% for SCC with sensitivities above 90%.

Machine learning radiomics models showed accuracy of 75-87%, with the best results combining radiomics features with expert-guided tumor delineation or multiphasic contrast-enhanced CT. Linning et al.'s radiomics model using venous-phase CT achieved an AUC of 0.864 for differentiating adenocarcinoma from SCC. Importantly, in one study (Bashir et al.), semantic features interpreted by radiologists outperformed purely computational radiomics on external validation (AUC 0.82 vs. 0.52), highlighting the value of domain expertise.

Combined approaches showed particular promise: Tang et al.'s ensemble combining intratumoral and peritumoral radiomics features with five classifiers achieved an AUC of 0.87 in training and 0.78 in independent external testing. The Marentakis et al. LSTM+Inception deep learning model outperformed expert radiologists on a histology classification task, reaching 74% accuracy on a dataset where traditional methods struggled.

Saad et al.'s hybrid Genetic Algorithm-SVM approach achieved 96.2% subtyping accuracy and demonstrated survival stratification value, with a concordance correlation coefficient of 0.99 compared to pathological diagnosis - suggesting AI classification could have genuine prognostic utility beyond just histologic labeling.

TL;DR: Deep learning systems achieved up to 97% accuracy and outperformed radiomics-only models in most comparisons, while combined intratumoral-peritumoral radiomics with ensemble learning showed strong generalization to external test sets.
Page 8
Where AI Models Fail: Challenging Imaging Phenotypes

Consistent patterns of misclassification emerged across multiple studies, revealing specific imaging scenarios where AI models struggle. SCC was more frequently misclassified than ADC, with SCC false-negative rates of 31-37% in several studies compared to ADC false-negative rates of only 8-19%. This asymmetric difficulty reflects the greater phenotypic variability of SCC appearance on CT.

The most common failure modes were associated with: overlapping tumor density between SCC and ADC, central necrosis and cavitation (air-containing holes within tumors) which can appear in both subtypes, central airway location (where SCC is more common but not exclusive), absence of ground-glass opacity (a fuzzy haziness more typical of ADC), and ambiguous tumor margin morphology.

Technical factors also contributed significantly to misclassification: small lesion size (5 mm or less), variability in scanner types and imaging protocols between training and test datasets, and differences in CT slice thickness and reconstruction algorithms all degraded model performance. SCC consistently emerged as the weakest class in multiclass models, with lower F1 scores and higher misclassification rates than ADC or SCLC.

TL;DR: SCC is consistently harder to classify than ADC, with false-negative rates up to 37%, driven by phenotypic overlap, cavitation, central location, and scanner variability - revealing where future model improvements are most needed.
Pages 4-5
Imaging Protocol Variability and Reproducibility Challenges

A critical challenge identified across all 11 studies is the lack of standardization in CT acquisition protocols. Scanner manufacturers ranged from Siemens, GE, Philips, Toshiba, and United Imaging systems, each with different reconstruction algorithms that alter how tumor textures appear. Slice thickness varied widely from 0.7 mm to 5 mm - thinner slices provide more detail but increase noise, while thicker slices smooth fine textural features that radiomic models depend on.

Contrast agent use was inconsistent, with some studies using non-contrast CT, others using contrast with arterial and/or venous phase imaging. Since contrast changes how tumors appear (brightening well-vascularized regions), mixing contrast and non-contrast images in training data introduces confounding variation. Interestingly, Linning et al. found that the contrast phase (venous being best) affected performance but without large differences, suggesting radiomics features have some robustness to contrast timing.

Tumor segmentation methodology varied substantially across studies, ranging from single-reader manual delineation to multi-observer consensus contouring and automated approaches. Haga et al. specifically showed that features selected based on interobserver delineation stability (AUC 0.725) performed somewhat worse than features selected independently (AUC 0.757), but both were more reliable than features that varied between annotators.

TL;DR: Across the 11 studies, wide variation in scanner types, slice thicknesses, contrast protocols, and segmentation methods severely limits cross-study comparability and real-world deployment of AI models.
Pages 11-12
Architecture Comparison: What Works Best

Among deep learning architectures, no single approach dominated across all settings. 3D CNNs (like ProNet) performed well on large datasets by processing the full tumor volume, achieving AUC 0.84. Capsule Networks (CapsNet) showed particular strength on smaller datasets, achieving 81.3% accuracy by preserving spatial relationships between features better than standard CNNs. VGG16-based transfer learning worked effectively for histologic subtype classification on moderate-sized datasets.

The LSTM+Inception combination outperformed expert radiologists in one study by capturing temporal relationships between sequential CT slices in addition to spatial features within each slice - effectively processing the CT volume as a sequence of images and allowing the model to learn three-dimensional context without the full computational demands of a 3D CNN.

For radiomics-based models, ensemble classifiers combining multiple algorithms consistently outperformed single classifiers. The Tang et al. ensemble combining five machine learning methods (QDA, SVM with two kernels, Random Forest) achieved the best generalization to external test data, suggesting that model diversity compensates for individual classifier weaknesses and produces more robust predictions.

TL;DR: 3D CNNs and Capsule Networks led among deep learning approaches, while ensemble classifiers combining multiple algorithms showed best generalization for radiomics-based models - with dataset size and diversity being critical factors in all cases.
Page 13
Future Directions and Clinical Translation Requirements

All included studies received RQS scores below 50%, revealing systematic gaps in methodological rigor: absent phantom studies for feature stability testing, no prospective validation designs, lack of cost-effectiveness analyses, missing calibration statistics, and insufficient reporting of inter-observer segmentation agreement. Addressing these gaps is essential before any of these AI tools could be adopted in routine clinical practice.

External validation in independent patient cohorts remains sparse. Only four of the 11 studies explicitly reported external validation results, and in these, performance consistently dropped compared to internal training performance - highlighting the overfitting risk when models are evaluated only on the data they were trained on. Future studies must include rigorous external validation to demonstrate true generalizability.

The most promising path forward combines AI imaging analysis with clinical and pathological context. The systematic review confirmed that combining radiomic features with expert-derived semantic descriptors, clinical variables, and pathological data consistently outperformed any single data source. Building multimodal AI systems that integrate these complementary information streams - rather than relying on imaging alone - will be essential for achieving clinically acceptable diagnostic accuracy in practice.

Standardization of CT acquisition protocols, feature extraction pipelines, and reporting standards across institutions is the foundational requirement for making AI-based lung cancer subtyping a reliable clinical tool. The field would benefit enormously from prospective multicenter studies that pre-specify imaging protocols, use standardized segmentation approaches, and validate models in genuinely independent patient populations before any clinical deployment.

TL;DR: While AI demonstrates strong potential for non-invasive lung cancer subtype classification, all reviewed studies scored below 50% on quality metrics, and extensive external validation, protocol standardization, and multimodal integration are needed before clinical use.
Citation: Open Access, 2026. Available at: PMC12896415.