The challenge of tumor heterogeneity. Non-small cell lung cancer (NSCLC) accounts for 85% of all lung cancer cases, yet its low survival rate persists because first-line therapies frequently fail due to insufficient understanding of molecular-level tumor heterogeneity. Different tumors behave differently at the molecular level, and current imaging and clinical tools cannot capture this complexity.
Why RNA regulation matters. Cancer is not driven only by protein-coding genes. Long non-coding RNAs (lncRNAs) and microRNAs (miRNAs) form complex regulatory networks that control gene expression without directly encoding proteins. In many cancers, these non-coding RNAs act as molecular sponges or switches that alter the expression of key cancer-driving genes.
Machine learning as a multi-omics integrator. High-throughput sequencing generates massive datasets spanning mRNA, miRNA, lncRNA, and protein expression simultaneously. Machine learning algorithms can process these large multi-omics datasets to identify which molecular features consistently distinguish cancer from normal tissue and which combinations of genes best predict disease behavior.
Study goal. This study integrated mRNA, lncRNA, miRNA, and protein expression data from publicly available NSCLC datasets, then applied machine learning and network analysis to identify diagnostic biomarkers, key regulatory axes, and candidate drug targets in NSCLC.
Five publicly available datasets. Gene expression data were obtained from the Gene Expression Omnibus (GEO) database. Three datasets were used for differential expression analysis: GSE27262 and GSE116959 for mRNA and lncRNA profiles, and GSE51853 for miRNA profiles. GSE75037 was used for machine learning-based feature selection, and GSE118370 served as the independent validation dataset. Together, these datasets encompassed hundreds of NSCLC tumor and normal tissue samples across multiple platforms.
Identifying differentially expressed genes. The limma package in R was used to compare cancer versus normal expression across datasets. For mRNAs and lncRNAs, genes were considered differentially expressed if they showed at least a 2-fold change with an adjusted p-value below 0.05. For miRNAs, a slightly reduced fold-change threshold of 1.74-fold was used because miRNAs naturally exhibit smaller expression changes than mRNAs. A total of 326 genes were consistently upregulated and 679 consistently downregulated across both mRNA datasets.
WGCNA to find co-expressed gene modules. Weighted Gene Co-expression Network Analysis (WGCNA) groups genes with similar expression patterns across patients into modules. Rather than looking at each gene individually, WGCNA identifies clusters of genes that rise and fall together, which are more likely to share biological functions. The modules most strongly correlated with the tumor phenotype were identified in both datasets, and their overlapping genes were extracted, yielding 227 shared module genes.
Converging on 57 candidates. The 227 shared WGCNA module genes were intersected with the 1,005 consistently differentially expressed genes, producing a final set of 57 high-confidence candidate genes that are both differentially expressed in NSCLC and co-regulated with cancer-associated expression patterns.
Building a protein-protein interaction network. The 57 candidate genes were uploaded to the STRING database to build a protein-protein interaction (PPI) network with 57 nodes and 478 interactions. This network visualizes how proteins physically interact with or influence each other, revealing which genes are most central to the overall molecular machinery of NSCLC.
Finding the most connected hub genes. Two complementary network analysis tools were used. The MCODE plugin identified the most densely connected submodule containing 22 genes including TOP2A, CDK1, AURKA, BUB1B, CENPF, and TPX2. The CytoHubba plugin scored all genes on multiple centrality metrics including degree, betweenness, closeness, stress, and maximal clique centrality, then ranked the top 10: CDK1, TOP2A, CCNB1, CCNB2, AURKA, TPX2, MAD2L1, BUB1B, UBE2C, and CENPF.
Three machine learning algorithms applied in parallel. LASSO regression, Random Forest, and Support Vector Machine with Recursive Feature Elimination (SVM-RFE) were each independently applied to the GSE75037 dataset to select the most diagnostically informative genes. LASSO selected 6 genes. Random Forest and SVM-RFE each selected 10 genes. The intersection across all three methods identified 6 genes that all algorithms agreed on: AURKA, CDK1, TPX2, TOP2A, BUB1B, and CENPF.
Validation with AUC over 0.8. These six biomarkers were then tested on the independent GSE118370 validation dataset. ROC curve analysis showed that CDK1, TOP2A, TPX2, CENPF, and AURKA each achieved an AUC greater than 0.8, confirming their strong diagnostic capability for distinguishing NSCLC from normal lung tissue in a held-out dataset not used for feature selection.
What is a ceRNA network. A competing endogenous RNA (ceRNA) network describes how lncRNAs can act as molecular sponges, soaking up miRNAs and preventing them from silencing their target mRNAs. In this sponge model, if a lncRNA is overexpressed, it sequesters miRNAs, which frees up mRNAs to be translated into more protein, potentially driving cancer progression.
Building the NSCLC ceRNA network. miRcode and DIANA-LncBase databases were used to identify which differentially expressed lncRNAs interact with which differentially expressed miRNAs. Then miRDB and TargetScan were used to identify which of those miRNAs target differentially expressed mRNAs. Interactions were retained only when expression patterns were inverse (lncRNA high when miRNA low, miRNA high when mRNA low), consistent with the sponge mechanism. This produced a final ceRNA network of 6 lncRNAs, 10 miRNAs, and 30 mRNAs.
CDK1 as the hub of the network. Among the six machine learning-identified biomarkers, only CDK1 appeared within the ceRNA network. The interaction pathway connecting to CDK1 was identified as: PVT1 (the oncogenic lncRNA) sequesters miR-143-3p (a tumor-suppressor microRNA), which would otherwise silence CDK1 (a cell cycle kinase). When PVT1 is overexpressed, miR-143-3p is tied up, CDK1 is derepressed, and cell cycle progression accelerates.
Transcription factors governing the axis. NetworkAnalyst and JASPAR database analysis identified FOXC1, YY1, and GATA2 as the three transcription factors with the highest connectivity to genes in the ceRNA network, specifically as regulators of CDK1. These three transcription factors have also been identified as key regulators of differentially expressed genes in other NSCLC bioinformatics studies, suggesting they orchestrate a broader transcriptional program in this cancer.
Expression confirmed in TCGA-LUAD. Using the UALCAN database, the expression of all axis components was confirmed to be significantly altered between tumor and normal samples in The Cancer Genome Atlas lung adenocarcinoma (TCGA-LUAD) dataset. PVT1, miR-143-3p, CDK1, FOXC1, YY1, and GATA2 all showed statistically significant differences in expression, validating the relevance of these findings in a large independent clinical dataset.
Survival associations are clinically meaningful. Kaplan-Meier survival analysis using the KM Plotter database showed that CDK1, PVT1, GATA2, and YY1 all have statistically significant associations with overall survival in NSCLC patients (p less than 0.05). High expression of these genes correlates with worse survival, reinforcing their role as both diagnostic biomarkers and potential prognostic factors.
Strong molecular interaction scores for the axis. The interaction between PVT1 and miR-143-3p was characterized by a TDMDScore of 0.9593, indicating a strong potential for target-directed miRNA degradation. The interaction between miR-143-3p and CDK1 had an even higher TDMDScore of 1.0546. Both interactions were supported by Argonaute CLIP-seq experimental data, providing biochemical evidence beyond computational prediction.
Pan-cancer relevance. The PVT1/miR-143-3p interaction was detected across eight different cancer types in TCGA data, and the miR-143-3p/CDK1 interaction was observed in seven cancer types. This broad pan-cancer relevance suggests that the regulatory axis is not unique to NSCLC and may represent a more general cancer biology mechanism, increasing confidence in its functional importance.
225 candidate therapeutic compounds identified. The Drug-Gene Interaction Database (DGIdb) was used to screen the six machine-learning-identified biomarkers against known drug-gene interactions. A total of 225 candidate compounds were identified: 50 targeting CDK1 (including Dinaciclib, which has been shown to induce anaphase catastrophe in lung cancer cells), 58 targeting AURKA (including Alisertib, which has been evaluated in a phase 2 clinical trial in NSCLC), and 117 targeting TOP2A (including Amonafide and Aldoxorubicin).
Multiple intervention points in the PVT1 axis. The ceRNA axis offers at least three therapeutic entry points. PVT1 could be silenced using antisense oligonucleotides, reducing its sponging of miR-143-3p. Alternatively, synthetic miR-143-3p mimics could restore the miRNA's tumor-suppressive function, reducing CDK1 expression. CDK1 itself can be directly targeted by small molecule inhibitors like Dinaciclib. Combinations of these approaches could have synergistic effects.
Endocrine resistance pathway connection. KEGG pathway analysis of the hub axis genes revealed that the most significantly enriched pathway is the endocrine resistance pathway. This unexpected finding suggests that the PVT1/miR-143-3p/CDK1 axis may contribute not only to cell cycle dysregulation but also to resistance mechanisms against endocrine therapies, a connection that warrants further investigation.
Limitations and future work. All findings in this study are computational and based on publicly available databases. The ceRNA axis is described as hypothetical and requires experimental validation in cell line and animal models before its mechanistic role can be confirmed. Clinical studies will also be needed to validate the diagnostic utility of the six biomarkers and assess the therapeutic efficacy of the predicted drug candidates in NSCLC patients.
Six key biomarkers consistently identified. The integration of network analysis and three independent machine learning algorithms converged on CDK1, TOP2A, AURKA, TPX2, BUB1B, and CENPF as the most reliable biomarkers in NSCLC. These genes are all involved in cell cycle regulation, consistent with the well-established hallmark of cancer involving uncontrolled proliferation and genomic instability.
A new systems-level regulatory model. While individual components of the PVT1/miR-143-3p/CDK1 axis have been studied separately, this study is the first to combine them into a single ceRNA regulatory network that links non-coding RNA biology with a central cell cycle kinase. This systems-level perspective connects upstream transcriptional regulators (FOXC1, YY1, GATA2) to non-coding RNA regulation to downstream cell cycle control, offering a more complete model of NSCLC molecular regulation.
Clinical translation potential. The combination of diagnostic biomarkers (each achieving AUC above 0.8) with drug target identification and survival associations provides a pipeline from molecular discovery to potential clinical application. The validated survival associations of CDK1, PVT1, GATA2, and YY1 suggest these could serve as prognostic markers in addition to diagnostic ones.