The central question. CT and FDG PET/CT are the two main imaging tools used in lung cancer care -- CT for screening and structural assessment, PET/CT for staging and treatment planning. AI models have been built on each, and on combinations of both. This systematic review asked: in direct head-to-head comparisons within the same patient cohorts, which modality combination performs best, and for what clinical tasks?
Evidence gap being addressed. Most existing AI reviews are broad narratives that fail to compare modalities directly within the same cohorts using the same patient groups and splits. Without same-cohort head-to-head comparisons, it is impossible to know whether one modality genuinely outperforms another or whether differences reflect patient selection bias.
Scope and methods. This PRISMA 2020-compliant systematic review searched PubMed, Scopus, IEEE Xplore, and Google Scholar for English-language studies published between January 2019 and September 2025. From 2,417 records identified, 31 studies met inclusion criteria: 20 primary modeling studies and 11 narrative reviews. Risk of bias was assessed using PROBAST for prediction studies and SANRA for reviews.
Core finding. No single modality or AI approach dominates across all clinical contexts. CT-centered models are strongest for population screening, PET/CT fusion excels for nodal staging and prognosis, and adding clinical variables provides consistent incremental benefit across all settings. The best choice depends fundamentally on the clinical task.
Strict inclusion standards. To ensure only meaningful comparisons were included, studies needed to perform explicit same-cohort, same-split comparisons between CT-only, PET/CT-only, and combined multimodal AI models. Tasks had to cover screening and nodule risk stratification, nodal metastasis or staging, or prognosis and recurrence prediction. Studies without a direct head-to-head comparison within the same cohort were excluded.
Database coverage. Searches combined controlled vocabulary (MeSH terms) and free-text terms across four databases. The search covered lung cancer, CT, PET/CT, comparison, multimodal fusion, screening, staging, prognosis, and survival. Queries were iteratively refined with an information specialist and adapted to each database's syntax. Manual search of Google Scholar's top 200 results per query supplemented database searching to capture in-press and early-access articles.
Quality assessment tools. Two validated instruments were applied. PROBAST assessed prediction model studies across four domains: participants, predictors, outcome, and analysis. SANRA evaluated narrative reviews on six dimensions including justification of importance, literature search quality, scientific reasoning, and presentation. All assessments were performed by multiple independent reviewers with discrepancies resolved by discussion and senior adjudication.
Risk of bias findings. Of 20 primary modeling studies, only one was judged overall low risk of bias. Twelve had some concerns, typically from limited calibration reporting, modest external validation, or small event-per-predictor ratios. Seven were high risk, driven by single-center internal splits, small samples with complex multimodal fusion, inappropriate metrics for time-to-event outcomes, or pronounced protocol heterogeneity. High-risk judgments clustered in the analysis domain.
Screening: CT dominates. In lung cancer screening cohorts using low-dose CT (LDCT), deep learning CT models consistently formed the strongest unimodal baseline. Adding tabular clinical variables to CT produced consistent but modest incremental improvements. Crucially, FDG PET/CT is not part of screening workflows and was not evaluated head-to-head against LDCT in any screening-specific cohort, so screening conclusions are necessarily CT-centered.
Nodal staging: PET/CT fusion wins decisively. A carefully controlled same-cohort study directly compared CT radiomics, PET radiomics, PET/CT radiomics, a clinical model, and an integrated PET/CT radiomics-clinical model for preoperative thoracic lymph node metastasis prediction. The integrated PET/CT radiomics-clinical model achieved the highest AUC (approximately 0.94) with favorable calibration and superior decision-curve net benefit, outperforming all unimodal approaches in the same cohort with identical splits.
Prognosis and recurrence: fused PET/CT outperforms individual modalities. Multiple head-to-head same-cohort studies evaluating overall survival and postoperative recurrence risk consistently found PET/CT radiomics to outperform PET-only and CT-only models. In two-institution studies that trained on one site and tested on the other, the PET/CT fusion advantage held across cross-institutional validation. Adding clinical covariates further improved performance beyond the fused imaging model.
Context dependency confirmed. Applying models outside their development context caused notable performance drops. Screening-optimized single-timepoint CT models degraded substantially when applied to incidentally detected or biopsy-selected cohorts. This spectrum bias confirms that modality and modeling choices must match the specific clinical use case rather than being assumed to generalize across the care continuum.
Screening: intermediate fusion of CT with tabular data. For screening, the most effective approach was intermediate co-learning that jointly trains on whole-scan CT and clinical data simultaneously, rather than combining them as separate late-stage inputs. This approach outperformed both CT-only and clinical-only models and generalized to external geographic test sets. A multimodal model that added blood biomarkers alongside CT and clinical data further improved discrimination, and critically, was designed to handle missing biomarkers at inference -- a practically important property in screening programs where ancillary data are often unavailable.
Nodal staging: feature-level PET/CT fusion with clinical variables. The strongest head-to-head evidence for nodal staging came from feature-level fusion of PET and CT radiomic features, further augmented by clinical variables. This approach achieved the best combination of discrimination, calibration, and clinical utility (decision-curve analysis) in the same cohort with identical splits. Feature-level fusion at this scale captures the complementary metabolic and morphologic information in spatially co-registered PET/CT images.
Prognosis: both image-level and feature-level PET/CT fusion work. For survival and recurrence prediction, two PET/CT fusion strategies consistently outperformed single-modality models across institutions. Feature-level fusion concatenated modality-specific radiomic features, while image-level fusion constructed a wavelet-blended PET/CT image before extraction. Both approaches held their advantages under cross-site evaluation, suggesting the benefit reflects complementary biology rather than site-specific confounding. Clinical covariates provided further incremental benefit on top of the fused imaging.
When PET is unavailable: CT with clinical data still adds value. In settings where PET/CT is not accessible, combining deep CT features with clinicopathologic variables using feature-level or late fusion produced small but statistically significant discrimination gains and better decision-curve net benefit than either source alone. This finding was supported by external temporal and geographic validation, making it clinically actionable for CT-only workflows.
Clinical variables: consistent benefit across all settings. Adding clinical variables to imaging models improved discrimination in every clinical context studied. For screening, joint co-learning of CT with clinical data elements improved performance internally on the National Lung Screening Trial (NLST) and externally in a separate cohort. For nodal staging, adding clinical variables to PET/CT radiomics yielded the highest test AUC with favorable calibration and decision-curve benefit. For postoperative progression, adding clinicopathologic variables to CT deep features produced a significant improvement and better clinical utility.
Blood biomarkers: promising but requires missingness handling. A multi-path network integrating CT, clinical variables, and serum biomarkers significantly outperformed any single modality. Critically, this architecture was designed to handle missing biomarkers during both training and inference, and training on incomplete cases actually improved discriminatory power compared to using only complete cases. This missingness-aware design is rare in the literature but represents best practice for real-world deployment.
Histopathology and genomics: promising but limited evidence. Late fusion of CT radiomics with whole-slide imaging pathomics and clinical variables outperformed each unimodal model in a radiotherapy-focused study. Radiogenomics studies combining CT with RNA sequencing reported concordance index improvements of up to approximately 10% over either modality alone. However, these findings rest on small single-center cohorts without external validation and should be considered hypothesis-generating rather than practice-changing.
Missing modality handling: a critical gap. Most multimodal studies assumed complete availability of all data types at both training and inference. Only one study explicitly modeled missing modalities during training and inference. This gap severely limits real-world applicability because in clinical practice, blood biomarkers, genomic data, and even complete clinical records are frequently unavailable for individual patients.
External validation is essential for credibility. Studies incorporating external, temporal, or cross-site validation consistently reported more conservative but more credible performance estimates. Single-center studies with internal cross-validation frequently produced higher metrics that did not reliably transport across settings. Cross-site train/test swaps for PET/CT survival modeling confirmed that fusion advantages over single modalities persisted under distribution shift, providing the strongest evidence that the benefit reflects true complementary biology.
Spectrum effects across cohort types. A systematic evaluation that re-implemented diverse risk models across screening-detected, incidentally detected, and biopsy-selected cohorts found that performance rankings shifted substantially across these populations. Models optimized for screening degraded substantially in biopsy-enriched cohorts, and no single approach dominated across all three contexts. This spectrum dependency must be accounted for when selecting AI tools for specific clinical workflows.
Harmonization gaps undermine generalizability. Only approximately one-third of studies explicitly reported IBSI-compliant CT radiomics pipelines or PET/CT standardization steps, and fewer than 15% applied formal statistical harmonization such as ComBat. A multicenter delta-PET/CT study directly demonstrated that protocol heterogeneity across scanners can completely erase apparent prognostic signal, warning that positive internal results may vanish under cross-site testing without harmonization. PET/CT reviews consistently called for SUV normalization and reconstruction standardization as prerequisites for reproducibility.
Metric reporting: discrimination without utility is insufficient. Approximately 80% of studies reported appropriate discriminative metrics, but fewer than 25% supplemented these with calibration assessment, and only 10 to 15% reported decision-curve analysis. For clinical adoption, knowing whether a model improves discrimination is necessary but not sufficient -- calibration (are predicted probabilities accurate?) and net benefit (does using the model improve patient outcomes compared to treat-all or treat-none strategies?) are equally important.
Task-specific modality selection. The review synthesizes evidence into a clear framework: population screening benefits most from CT with clinical covariates; incidentally detected nodules benefit from CT plus or minus PET/CT with clinical data; preoperative nodal staging performs best with PET/CT plus clinical variables using feature-level fusion; prognosis and recurrence prediction favor PET/CT with image-level or feature-level fusion. No single strategy works best across all tasks.
Algorithmic complexity is not the primary driver of performance. Differences in reported performance across studies were driven primarily by clinical context, cohort selection, imaging acquisition protocols, outcome definitions, and validation design rather than by which specific AI algorithm was used. This finding emphasizes that data quality, study design, and validation rigor matter more than algorithm sophistication for translating AI research into reliable clinical tools.
Radiogenomics and pathomics: handle with caution. Integrating CT with tumor gene expression or whole-slide image pathology provides added prognostic signal and is biologically compelling. However, these approaches consistently rested on small samples, lacked external validation, and omitted calibration and utility analyses. They represent promising research directions rather than tools ready for clinical deployment.
Roadmap for future research. The field needs more multicenter prospective studies with pre-specified same-cohort head-to-head modality comparisons, standardized imaging acquisition and reconstruction protocols (especially for PET/CT), explicit missing-modality handling in multimodal architectures, and comprehensive reporting of calibration and decision-curve analysis alongside discrimination metrics. These improvements are prerequisites for evidence-based modality selection in clinical AI deployment.