Non-small cell lung cancer is a leading cancer killer. Non-small cell lung cancer (NSCLC) accounts for roughly 80% of all lung cancer cases, encompassing adenocarcinoma, squamous cell carcinoma, and large cell carcinoma. Despite advances in surgery, chemotherapy, and targeted therapies, outcomes remain poor due to late-stage diagnosis and drug resistance - making improved understanding of the disease's causes critically important.
Air pollution is a known lung cancer risk factor. Research has established that environmental air pollutants increase the risk of NSCLC and may accelerate its progression. Common pollutants studied include gases like carbon monoxide (CO), nitric oxide (NO), nitrogen dioxide (NO2), and sulfur dioxide (SO2), as well as polycyclic aromatic hydrocarbons like benzo[a]pyrene (BaP) and 3-methylcholanthrene (3-MC) produced by combustion processes.
The molecular mechanisms remain unclear. While the epidemiological link between air pollution and lung cancer is well established, the specific molecular pathways through which pollutants promote cancer development are less understood. Identifying which genes are affected by air pollutants, and how those same genes relate to lung cancer biology, could reveal both new biomarkers and therapeutic targets.
An integrative computational approach. This study combines three complementary computational methods: transcriptomic analysis of gene expression patterns in NSCLC patients, machine learning to identify the most diagnostically relevant genes from large datasets, and molecular docking simulations to examine how air pollutant molecules physically interact with key proteins at the atomic level.
Six gene expression datasets from public repositories. The study curated six transcriptomic datasets from the NCBI Gene Expression Omnibus (GEO) database, capturing gene activity levels measured in NSCLC tumor tissue compared to normal lung tissue. Three datasets formed the training cohort and three formed an independent validation cohort, enabling proper testing of whether findings generalize beyond the initial data.
Correcting for batch effects. A major challenge in combining data from different experiments is 'batch effects' - systematic differences in measurements caused by different labs, equipment, or protocols rather than biology. The researchers applied Surrogate Variable Analysis followed by ComBat harmonization to correct these biases, verified using principal component analysis to confirm that samples from different batches now clustered by biology rather than source.
Identifying air pollutant gene targets. The molecular targets of seven air pollutants were identified using three complementary databases: ChEMBL (which catalogs ligand-receptor interactions), SwissTarget Prediction (which predicts targets from chemical structure), and PharmMapper (which matches 3D molecular shapes to protein binding sites). Combining all three sources yielded 1,242 unique gene targets associated with the seven pollutants.
Weighting gene co-expression networks. Beyond simply identifying differentially expressed genes, the researchers used Weighted Gene Co-Expression Network Analysis (WGCNA) to find groups of genes that are expressed together in coordinated patterns. Genes that are tightly co-expressed and strongly correlated with disease status (the MEblue module in this study) are more likely to be functionally important in NSCLC biology.
Finding the overlap. By intersecting the 1,242 genes targeted by air pollutants with the 247 genes found to be significantly altered in NSCLC through differential expression and co-expression analysis, the researchers identified 30 genes that sit at the crossroads of pollution exposure and lung cancer biology. These 30 genes represent the most promising candidates for linking environmental exposure to cancer development.
Cell cycle pathways dominate. Pathway enrichment analysis of these 30 genes revealed that the most significantly affected biological pathways include cell cycle regulation, cellular senescence, and p53 signaling. These are among the most fundamental cancer-related pathways - cell cycle dysregulation allows cancer cells to divide uncontrollably, cellular senescence represents a defense that cancer evades, and p53 is the most commonly mutated tumor suppressor gene in human cancers.
Protein-protein interaction network. The 30 genes were analyzed for how their protein products interact with each other using the STRING database. This protein-protein interaction network reveals which genes are central hubs - proteins that connect to many others and therefore have outsized influence on cellular biology. Highly connected hub genes in cancer networks are often especially relevant to disease progression and potentially to treatment.
Toxicological assessment of the seven pollutants. All seven air pollutants studied were assessed for carcinogenic potential using two independent computational platforms (ADMETLAB 3.0 and ProTox3). All seven were classified as carcinogenic by at least one platform. The polycyclic aromatic hydrocarbons (BaA, BaP, and 3-MC) showed the highest carcinogenic scores, consistent with their known ability to directly damage DNA and promote mutations.
A comprehensive machine learning comparison. To identify the most diagnostically relevant genes from the 30 candidates, the researchers applied 12 different machine learning algorithms - including LASSO regression, Random Forest, Support Vector Machine, XGBoost, Gradient Boosting, and others - generating 130 different predictive models in total. This exhaustive approach avoids bias toward any single method and identifies genes that are robustly important across multiple analytical frameworks.
Seven core genes emerged. The best-performing combination - LASSO regularization plus Random Forest feature selection - identified seven genes as the most informative diagnostic markers: CKS1B (involved in cell cycle regulation), GAPDH (a metabolic enzyme often dysregulated in cancer), TYMS (thymidylate synthase, involved in DNA synthesis), AURKA (Aurora kinase A, promotes cell division), CCNE1 (cyclin E1, drives cell cycle progression), PARP1 (involved in DNA repair), and MGLL (monoacylglycerol lipase, involved in lipid metabolism).
High diagnostic performance. ROC curve analysis showed that all seven core genes individually achieved area under the curve (AUC) values exceeding 0.95 in both training and validation datasets. This means the gene expression profile of these seven markers can distinguish NSCLC from normal tissue with over 95% discriminative accuracy - a level of performance that indicates genuine biological relevance rather than statistical artifact.
SHAP analysis reveals gene contributions. To interpret why the machine learning model makes its predictions, SHAP (SHapley Additive exPlanations) analysis was used to quantify each gene's individual contribution to diagnostic predictions. CKS1B and GAPDH emerged as the most significant predictors. Importantly, SHAP also revealed a non-linear inverse relationship between MGLL and CKS1B expression, suggesting these genes interact in complex ways during cancer development.
The immune landscape in NSCLC. Using the CIBERSORT computational algorithm, the researchers quantified the composition of 22 different immune cell types within both NSCLC tumor tissue and normal lung tissue. This type of immune deconvolution analysis infers the proportions of different immune cells from bulk gene expression data without requiring direct cell counting or tissue sampling of immune cells.
Significant immune differences in cancer. NSCLC tissues showed significantly different immune cell infiltration compared to normal lung, including altered levels of B cells (naive and memory), plasma cells, CD4+ T cell subsets, regulatory T cells, NK cells, monocytes, M1 macrophages, dendritic cells, mast cells, eosinophils, and neutrophils. This broad immune reshaping reflects how tumors remodel their microenvironment to evade immune destruction.
Core genes correlate with specific immune cells. The seven core air-pollutant-associated genes showed significant correlations with distinct immune cell populations within the tumor microenvironment. This suggests that the molecular alterations driven by air pollutant exposure may not only promote cancer cell growth directly but also reshape the immune environment in ways that reduce the body's ability to control tumor growth.
Implications for immunotherapy. Understanding how environmentally driven gene expression changes alter immune cell composition could help explain why some NSCLC patients respond better to immunotherapy than others. Patients with high environmental pollutant exposure may have distinctive immune microenvironment profiles that predict their response to treatments like PD-1/PD-L1 checkpoint inhibitors.
Molecular docking simulates physical binding. Molecular docking is a computational technique that predicts how a small molecule (here, an air pollutant) fits into the three-dimensional structure of a protein - similar to how a key fits a lock. By identifying the binding orientation with the lowest energy (most stable configuration), researchers can assess whether a pollutant could physically interact with a cancer-related protein and potentially disrupt its normal function.
All seven core genes can bind air pollutants. Docking simulations demonstrated that all seven core genes' protein products are capable of binding spontaneously to at least one of the seven air pollutants studied. The six pollutant-protein pairs with the strongest binding affinities were selected for detailed analysis: 3-MC binding GAPDH and MGLL, BaA binding AURKA, CCNE1, and PARP1, and BaP binding GAPDH.
Key interaction types identified. The strongest binding interactions were primarily hydrophobic (involving non-polar regions of the protein and pollutant), with van der Waals forces playing the dominant role across all six complexes. Specific amino acid residues that anchor each pollutant molecule within the protein's binding site were identified - for example, Leu-203 and Ala-238 on GAPDH form key interactions with 3-MC, while His-279 and Leu-194 on MGLL do the same.
Molecular dynamics confirms stability. To verify that the docking results represent stable, physiologically meaningful interactions rather than momentary configurations, 100-nanosecond molecular dynamics simulations were run for each complex. Analysis of structural metrics including root mean square deviation (RMSD) and binding free energy confirmed that all six pollutant-protein complexes remain stably bound under simulated physiological conditions, with van der Waals interactions as the primary stabilizing force.
Environmental factors shape cancer biology at the molecular level. This study provides computational evidence that specific air pollutants can physically interact with proteins encoded by genes known to drive NSCLC, and that exposure to these pollutants is associated with distinctive gene expression patterns in lung tumors. Together, these findings support a molecular mechanism through which air quality contributes to lung cancer risk beyond simply inhaling carcinogens.
The seven genes have known cancer relevance. Each of the seven identified core genes has prior evidence linking it to cancer biology. AURKA promotes abnormal cell division; CCNE1 drives uncontrolled cell cycle entry; PARP1 is involved in DNA repair pathways already targeted by approved cancer drugs; TYMS is the target of the chemotherapy drug fluorouracil. The identification of these genes through an air pollutant lens adds new context to understanding how environmental exposures engage already-known cancer mechanisms.
Potential for new biomarker development. With AUC values exceeding 0.95, the seven-gene signature shows genuine promise as a diagnostic tool. If validated in clinical samples, a blood or tissue test based on these genes could help identify NSCLC patients earlier or stratify them by environmental exposure history - potentially informing personalized treatment strategies that account for the molecular consequences of pollution exposure.
Limitations and next steps. The study is entirely computational - no laboratory experiments or patient samples were used to directly validate the identified genes. The datasets used are relatively small by modern standards, and causal inference (that pollution definitively causes these gene changes) cannot be made from this type of analysis alone. Future work must include in vitro cell experiments, animal models, and clinical cohort studies to confirm whether these computationally identified interactions translate to real-world cancer biology and clinical outcomes.
An integrative approach to environmental oncology. This study demonstrates the power of combining network toxicology, transcriptomics, machine learning, and structural biology to move beyond epidemiological observations toward molecular explanations of how air pollution promotes cancer. Rather than simply showing that pollution correlates with lung cancer risk, it identifies specific proteins, pathways, and interaction mechanisms as targets for further investigation.
Public health implications. If confirmed, these findings reinforce the scientific basis for stricter air quality regulations as a cancer prevention strategy. They also highlight that environmental pollution is not merely a background risk but actively engages cancer-driving molecular pathways - elevating air quality as a target not just for respiratory disease prevention but for cancer prevention more broadly.
Pathway to precision prevention. The identification of specific pollutant-gene interactions opens the possibility of developing 'exposure-aware' cancer risk models that incorporate both genetic predisposition and environmental exposure history. Such models could identify individuals at highest risk of developing pollution-associated NSCLC based on both their genetic profile and their residential air quality history, enabling targeted screening programs.