Cellular senescence is a state in which cells permanently stop dividing but remain metabolically active. While originally understood as a tumor-suppressing mechanism - stopping damaged cells from proliferating - recent research has revealed a more complex picture: senescent cells can also promote tumor growth by releasing inflammatory signals that alter the surrounding tissue environment.
In endometrial cancer, the role of cellular senescence-related genes in predicting outcomes has not been well characterized. If specific senescence genes could serve as reliable prognostic markers, they might identify patients at highest risk of recurrence who need more aggressive treatment, while sparing lower-risk patients from unnecessary treatment toxicity.
This study used data from 579 endometrial cancer patients in TCGA (The Cancer Genome Atlas), a large public database of cancer genomic and clinical data, to develop a machine learning-based prognostic model grounded in cellular senescence gene expression. Four different machine learning algorithms were tested and integrated to identify the most robust gene markers.
Starting from a list of 503 known cellular senescence-related genes, the researchers used consensus clustering - a statistical technique that groups patients based on similar gene expression patterns - to divide the 579 TCGA patients into two molecularly distinct groups. Cluster C1 contained 229 patients and cluster C2 contained 315 patients, with C2 showing significantly worse overall survival (OS) and progression-free survival (PFS), along with higher tumor cell content and lower immune cell infiltration.
To identify the specific senescence genes most responsible for these differences, the study applied WGCNA (Weighted Gene Co-expression Network Analysis), a method that groups genes that tend to be activated or suppressed together. WGCNA identified gene modules that were most strongly associated with the clinical outcome differences between the two clusters.
The candidate genes from WGCNA were then passed through four machine learning algorithms: GMM (Gaussian Mixture Model), SVM-RFE (Support Vector Machine with Recursive Feature Elimination), Random Forest, and XGBoost. Only genes selected by multiple algorithms were retained, yielding two final hub genes: CPEB1 and MYBL2. This consensus approach reduces the risk of identifying spurious markers that work in only one algorithm.
The two hub genes were combined into a simple risk score formula: Risk Score = (0.1279 x MYBL2 expression) + (0.0879 x CPEB1 expression). Patients with higher MYBL2 and CPEB1 expression were assigned higher risk scores. This linear formula is transparent and straightforward to apply once gene expression data is available, unlike black-box machine learning models.
MYBL2 is a transcription factor - a protein that turns other genes on or off - involved in cell cycle regulation. High MYBL2 expression promotes cell division and is associated with more aggressive cancer behavior. CPEB1 regulates mRNA translation (the process of reading genetic instructions to make proteins) and has been linked to tumor suppression; its inclusion in the high-risk group likely reflects dysregulation of normal cell control mechanisms.
Patients in the high-risk group had significantly worse overall survival compared with the low-risk group. The model's predictive accuracy, measured by AUC, was 0.624, 0.768, and 0.661 at 3 years across three independent cohorts - indicating moderate to good prognostic discrimination. The variation across cohorts reflects the challenge of generalizing genomic models across different patient populations and data collection methods.
A key clinical application of risk stratification is identifying which treatments are most likely to benefit patients in each risk group. The study analyzed drug sensitivity data and found that high-risk patients showed greater predicted sensitivity to several agents compared with low-risk patients.
Drugs predicted to be more effective in high-risk patients included Vincristine (a microtubule-targeting chemotherapy), BI-2536 (a PLK1 inhibitor that blocks a key cell cycle regulator), BMS-754807 (an IGF-1R/insulin receptor inhibitor), Bortezomib (a proteasome inhibitor used in blood cancers), and Daporinad (a NAMPT inhibitor targeting cancer cell metabolism). These differences in drug sensitivity are based on computational predictions from gene expression data rather than clinical trial evidence.
The finding that high-risk and low-risk patients show different drug sensitivity profiles has important implications: if validated, these results could guide personalized treatment selection. Rather than treating all high-risk patients with the same regimen, clinicians might choose agents based on a patient's molecular risk profile. However, these predictions require clinical validation before they can influence treatment decisions.
This study demonstrates that integrating cellular senescence biology with machine learning can generate clinically useful prognostic tools for endometrial cancer. The two-gene model is simple enough to be applied in clinical settings where RNA sequencing is available, and its biological basis in senescence pathways provides a coherent mechanistic rationale.
The link between high-risk senescence gene expression and lower immune infiltration in the tumor microenvironment is particularly interesting. Tumors with fewer immune cells are generally less responsive to immunotherapy - so risk group assignment might also predict which patients are candidates for immune checkpoint inhibitors versus other treatment approaches. This adds a layer of clinical utility beyond simple survival prediction.
Future validation in prospective clinical cohorts, and ultimately in clinical trials testing whether the model can guide treatment decisions, will be required before this tool can be implemented in routine care. The identification of MYBL2 and CPEB1 as key drivers also opens the door to developing novel drugs that specifically target these senescence-related pathways in high-risk endometrial cancer patients.