Prediction of final pathology depending on preoperative myometrial invasion and grade assessment in low-risk endometrial cancer patients: A Korean Gynecologic Oncology Group ancillary study

PLoS One 2024 AI 6 Explanations View Original
Original Paper (PDF)

Unable to display PDF. Download it here or view on PMC.

Plain-English Explanations
Pages 1-2
Fertility Preservation and the Preoperative Staging Problem in Low-Risk Endometrial Cancer

Hysterectomy is the primary treatment for endometrial cancer (EC), but progestin-based fertility-sparing treatment (FST) may be considered for reproductive-aged patients with low-risk disease. The NCCN Guidelines currently allow FST for patients with biopsy-proven grade 1 endometrioid adenocarcinoma with no myometrial invasion (MI) on MRI. Broader low-risk criteria - grade 1 or 2 adenocarcinoma with no MI or MI less than 1/2 - represent a larger population that might safely qualify for FST, given a lymph node metastasis risk of only 1.7-2.9%.

The critical obstacle to expanding FST eligibility is the poor concordance between preoperative assessment and final postoperative pathology. Prior studies show that preoperative MRI and endometrial biopsy frequently misclassify patients: the accuracy of MRI for MI assessment ranges from 54.8-86%, and matching rates between preoperative and postoperative MI are only 28.6-84.2% for no MI and 51.2-64.4% for MI less than 1/2. Grade concordance between biopsy and final pathology is 74.8-94.4% for grade 1 but only 43.8-58.8% for grade 2.

These discordance rates mean that some patients currently deemed low-risk preoperatively may actually have higher-stage or higher-grade disease at surgery. Conversely, some patients who could safely receive FST may be denied it due to imprecise preoperative staging. The clinical impact is significant: FST failure requiring eventual hysterectomy is associated with reduced quality of life and delayed fertility opportunities.

This study, an ancillary analysis (KGOG 2015S) of the Korean Gynecologic Oncology Group 2015 prospective multicenter cohort, aimed to quantify preoperative-postoperative concordance in low-risk EC patients and develop new machine learning prediction models to improve the prediction of final postoperative pathology from preoperative assessment variables.

TL;DR: Poor preoperative-to-postoperative concordance in low-risk endometrial cancer prevents reliable identification of FST candidates. This prospective multicenter study developed new machine learning models to predict final pathology from preoperative MRI and biopsy findings.
Pages 2-3
KGOG 2015 Cohort and Four-Group Stratification Design

The parent KGOG 2015 study enrolled 529 consecutive EC patients at 20 tertiary hospitals across Korea, Japan, and China between January 2012 and December 2014. All patients underwent preoperative MRI, endometrial biopsy, and serum CA-125 testing, followed by surgical staging including systematic pelvic and para-aortic lymphadenectomy. Staging used 2009 FIGO criteria based on final pathological findings.

For this ancillary study (KGOG 2015S), 251 eligible patients were selected who had preoperative MRI showing no MI or MI less than 1/2, and biopsy-proven endometrioid adenocarcinoma of grade 1 or 2. Patients were assigned to four groups: Group 1 (no MI, grade 1, n=106), Group 2 (no MI, grade 2, n=41), Group 3 (MI less than 1/2, grade 1, n=74), and Group 4 (MI less than 1/2, grade 2, n=30). Mean age was 52.8 years; 60.6% were postmenopausal; 86.1% underwent minimally invasive surgery.

Postoperative findings: Final pathology showed 90.4% of patients at FIGO stage IA, 4.8% stage IB, 2.4% stage II, and 2.4% stage IIIC. Grade distribution shifted postoperatively: 70.5% grade 1, 25.1% grade 2, and 2.0% grade 3. LVSI was present in 8.8%, cervical involvement in 2.4%, and pelvic lymph node metastasis in 2.4% of patients. The mean postoperative tumor sizes by MI category were 1.41 cm (no MI), 2.45 cm (MI less than 1/2), and 2.95 cm (MI 1/2 or more).

TL;DR: 251 low-risk EC patients from a 20-hospital prospective study were divided into four groups by preoperative MRI (no MI vs. MI less than 1/2) and biopsy grade (1 vs. 2). Most patients (90.4%) had FIGO stage IA at surgery, but significant upstaging occurred.
Pages 10-11
Poor Preoperative-Postoperative Concordance Across All Groups

Overall matching rates between preoperative group assignment and postoperative pathological classification were low: 43.4% in Group 1, 14.6% in Group 2, 60.8% in Group 3, and 43.3% in Group 4. Kappa statistics confirmed poor-to-fair agreement: k=0.304 for Group 1, k=0.145 for Group 2, k=0.281 for Group 3, and k=0.292 for Group 4. Group 2 (no MI, grade 2) was the most discordant group, with nearly 75% of patients upstaged at surgery.

Upstaging patterns: In Group 1, 56.6% of patients were upstaged (35.8% moved to Group 3, indicating occult MI less than 1/2 discovered at surgery). In Group 2, 75.6% were upstaged (26.8% to Group 3, 24.4% to Group 4, 24.4% to higher stages). In Group 3, 20.2% were upstaged. These upstaging rates indicate that preoperative MRI substantially underestimates myometrial invasion depth, particularly for no-MI cases.

MI concordance: MI matching rates were only 48.3% for no-MI cases (Group 1: 53.8%, Group 2: 34.1%) and 72.1% for MI less than 1/2 cases (Group 3: 74.3%, Group 4: 66.7%). Grade concordance was better: 84.4% for grade 1 and 56.3% for grade 2. The substantially lower concordance for MI compared to grade confirmed that preoperative MI measurement is the weaker predictor - a finding that directly shaped the NPM modeling strategy.

A total of 2.4% of patients had lymph node metastasis. Stage IIIC disease was found in 6 patients - all from Groups 1 and 2 (no MI preoperatively), highlighting that absence of MI on MRI does not guarantee low-stage disease. This reinforces the clinical need for better predictive tools before committing to fertility-sparing management or limited lymphadenectomy.

TL;DR: Preoperative-postoperative concordance was poor to fair across all groups (matching rates 14.6-60.8%). Group 2 (no MI, grade 2) was worst at 14.6% concordance, with 75.6% of patients upstaged at surgery. MI concordance (48-72%) was consistently worse than grade concordance.
Pages 4-9
NPM1 and NPM2: New Prediction Models Using Imputation and Label Smoothing

Two new prediction models (NPM1 and NPM2) were developed to predict postoperative pathological group from preoperative variables including age, menopause status, biopsy method, grade, MRI MI depth, tumor size, and serum CA-125. Both models use ensemble learning trained on class-balanced subsets with stratified K-fold cross-validation (K=5) to handle the significant class imbalance between the four study groups.

NPM1 - MI depth imputation as principal variable: The low concordance of MI depth between preoperative and postoperative assessments (average 58.2%) was addressed using iterative imputation based on Multivariate Imputation by Chained Equations (MICE). The preoperative MI depth values were treated as potentially unreliable and replaced with calibrated estimates derived from a model trained on seven correlated variables. During testing, the imputation model applied to the test set was then used as input to the ensemble classifier. NPM1 produced MSE=0.47 between calibrated and actual postoperative MI depths.

NPM2 - Removing MI depth from imputation with label smoothing: NPM2 identified a flaw in NPM1: using MI depth as a variable in its own imputation pulled calibrated values toward binary 0 or 1, losing the continuous variability of actual invasion depth within each category. NPM2 excluded MI depth from the imputation input variables, allowing imputed values to spread smoothly between 0 and 1, better capturing the real spectrum of invasion within each category. Label smoothing (converting hard labels to soft labels) was applied to reduce model overconfidence and improve generalization. NPM2 achieved MSE=0.31, significantly better than NPM1's 0.47.

Both NPMs outperformed conventional analysis and three baseline algorithms (logistic regression, XGBoost, SVM) on most metrics. Ensemble construction used the criterion: SM = argmax[(TP + TN) - (FP + FN)] to select the best-performing weak classifiers for combination. The ensemble models were compared against all four group classifications to identify where each approach provided the most clinically meaningful improvement in NPV - the primary outcome for ruling out high-risk disease.

TL;DR: NPM1 imputes preoperative MI depth using MICE to correct for MRI measurement error (MSE=0.47). NPM2 removes MI depth from the imputation variables and adds label smoothing, producing smoother MI calibration (MSE=0.31) and superior prediction performance in most groups.
Pages 11-12
Model Performance: NPV, Sensitivity, and AUC Across Groups

The new prediction models provided superior NPV and sensitivity compared to conventional analysis and baseline machine learning methods across all four groups. NPV is the most clinically relevant metric for FST candidate selection - it quantifies the probability that a patient classified as low-risk by the model is truly low-risk at surgery.

Group 1 (no MI, grade 1): Best NPV was 87.2% (NPM2), best sensitivity 71.6% (NPM1), best AUC 0.732 (NPM2). Logistic regression achieved the highest specificity (92.4%) and PPV (61.1%). For this largest group (106 patients), NPM2 provided the most reliable identification of patients who would truly remain group 1 at surgery.

Group 2 (no MI, grade 2): Best NPV was 97.6% (NPM1), best sensitivity 78.6% (NPM1), best AUC 0.656 (NPM2). Logistic regression achieved 100% specificity. Group 2 had the highest clinical stakes - with only 14.6% preoperative-postoperative matching, NPM1's 97.6% NPV represents the most dramatic improvement over conventional analysis, enabling safer identification of the rare Group 2 patients who truly remain low-risk postoperatively.

Groups 3 and 4: For Group 3 (MI less than 1/2, grade 1): best NPV 71.3% (NPM2), sensitivity 78.6% (NPM1), AUC 0.635 (conventional analysis, marginally). For Group 4 (MI less than 1/2, grade 2): best NPV 91.8% (NPM2), sensitivity 64.9% (NPM1), AUC 0.676% (NPM2). Overall, NPM2 was the best-performing model for NPV in 3 of the 4 groups, while NPM1 dominated sensitivity metrics. The two models complement each other: NPM2 for high-confidence negative prediction, NPM1 for minimizing missed cases.

TL;DR: NPM2 achieved the best NPV in 3 of 4 groups (87.2%, 71.3%, and 91.8%), while NPM1 achieved the best NPV of 97.6% in Group 2 and best sensitivity in most groups. Both models consistently outperformed logistic regression, XGBoost, and SVM on the most clinically relevant metrics.
Pages 12-14
Clinical Implications for Fertility-Sparing Treatment Decisions

The primary clinical application of this research is improving the identification of low-risk EC patients who are genuinely safe candidates for fertility-sparing treatment. The current finding that preoperative matching rates are as low as 14.6% in Group 2 patients means that the majority of these patients - assessed as low risk preoperatively - will be found to have higher-stage or higher-grade disease at surgery. This substantially limits FST candidacy under current criteria.

The high NPV values achieved by the NPMs (up to 97.6% in Group 2) are directly actionable: a clinician can use these models to identify which patients within each preoperative group have the lowest probability of being upstaged. Patients with model-predicted low risk can be counseled more confidently about FST, while those with higher predicted postoperative risk can be counseled toward definitive surgery earlier rather than after failed FST.

Lymph node metastasis occurred in 2.4% of the cohort, entirely in Groups 1 and 2 (no MI on preoperative MRI). This finding has implications for lymphadenectomy: even patients who appear low-risk on MRI are at non-negligible risk for nodal spread, particularly those with grade 2 biopsy. The prediction models may help stratify which no-MI patients warrant sentinel lymph node assessment or full lymphadenectomy versus observation.

Limitations: The study had no exclusion criteria beyond the four-group stratification, meaning heterogeneous subgroups exist within each group. Deep learning approaches were not explored due to the limited training dataset of 251 patients. The models require prospective validation in independent cohorts before clinical adoption. Future work should incorporate molecular markers (POLE mutation, mismatch repair status) that are now part of ESGO/ESMO risk stratification guidelines.

TL;DR: The NPMs enable more precise identification of FST candidates by predicting which preoperative low-risk patients will remain low-risk at surgery. High NPV values (87.2-97.6%) allow confident FST counseling; lymph node metastasis in no-MI patients underscores the need for careful surgical planning.
Citation: Open Access, 2024. Available at: PMC11210801.