Atypical endometrial hyperplasia (AEH) is a precancerous condition where the cells lining the uterus become abnormal. It is considered a significant risk factor for progressing to full endometrial cancer, and standard treatment guidelines recommend hysterectomy (surgical removal of the uterus) for most AEH patients.
However, not all patients can or want to undergo surgery. Some women with AEH or early-stage, low-grade endometrial cancer want to preserve their fertility. Others have serious health conditions that make surgery too risky. For these patients, hormonal treatment with progestins (a synthetic form of the hormone progesterone) is an alternative, as progestins can suppress abnormal endometrial cell growth.
The critical problem is that approximately 25% of patients do not respond to hormonal treatment - their AEH persists or progresses to cancer despite progestin therapy. Currently, there is no validated biomarker or test that can predict in advance which patients will respond. Clinicians must wait months to find out, during which time non-responders face ongoing cancer risk without effective treatment.
This study, conducted by researchers at the U.S. Food and Drug Administration (FDA) in collaboration with two academic medical centers, explored whether deep learning analysis of biopsy tissue slides could predict hormonal treatment response before treatment begins.
The study used whole slide images (WSIs) - high-resolution digital scans of complete tissue biopsy slides that can reach gigapixel (billions of pixels) in size. These were created by scanning formalin-fixed, paraffin-embedded endometrial biopsy slides using an Aperio AT2 scanner at 20x or 40x magnification.
The dataset included 112 patients from two clinical sites: Washington University School of Medicine in St. Louis (91 patients) and the University of Oklahoma Health Sciences Center (21 patients). All patients had been diagnosed with either AEH or FIGO grade 1 (low-grade) endometrial cancer, received progestin treatment, and had a follow-up biopsy 3-15 months later to determine response.
Patients were labeled as responders (follow-up biopsy showed no remaining atypia) or non-responders (follow-up biopsy still showed atypical hyperplasia or carcinoma). Of the 112 patients, 66 were responders and 46 were non-responders.
An expert board-certified pathologist annotated the diagnostic slides by outlining regions containing AEH or cancer tissue in the images. A second pathologist reviewed all annotations for consensus. These annotations defined which tissue areas the AI system would analyze - a step the authors call mixed supervision, because the pathologist fully supervises which regions are analyzable, while the AI training itself uses only the coarser binary responder/non-responder labels.
The researchers compared three different methods of converting image patches into numerical feature vectors that a classifier could use to predict treatment response. Each approach represents a different philosophy about how to extract meaningful information from tissue images.
The first approach was a convolutional autoencoder (CAE) - a self-supervised neural network trained on the study's own histology data. An autoencoder learns to compress an image into a compact representation (the latent vector) and then reconstruct the original image from that representation. The quality of reconstruction forces the network to learn the most important visual features. This approach requires no labels during the feature extraction training phase.
The second approach used a ResNet50 model pre-trained on ImageNet (a massive database of everyday photographs) to extract features through transfer learning. ResNet50 is a well-established architecture, but features learned from photographs of everyday objects may not be optimally suited to microscopic tissue images.
The third approach used radiomics feature extraction via PyRadiomics, a software package that calculates 94 predefined mathematical features from images - including texture measures, intensity statistics, and shape properties. Unlike deep learning, radiomics does not require training; it applies fixed mathematical formulas to extract features. After extraction by any of these three methods, a fully connected neural network was trained on the resulting feature vectors to make the binary responder/non-responder prediction.
The primary performance metric was AUROC (Area Under the Receiver Operating Characteristic Curve), which measures how well a model can distinguish between two groups (responders versus non-responders) across all possible decision thresholds. An AUROC of 1.0 is perfect; 0.5 is equivalent to random guessing.
The autoencoder model achieved an AUROC of 0.80 (95% confidence interval: 0.63-0.95) on the independent test set. This is a meaningful result for a feasibility study - it suggests the AI is capturing real biological signals in the tissue images that correlate with treatment response, substantially above chance.
In comparison, the ResNet50 model achieved only 0.60 AUROC and the radiomics approach achieved 0.67. The autoencoder's superiority may be because it was trained on endometrial histology images specifically, allowing it to learn tissue-relevant features, while ResNet50 features transferred from everyday photographs may not align well with the subtle pathological differences in biopsy tissue.
However, the difference in performance between the three models was not statistically significant (p-values of 0.22 and 0.68), due to the small test set of only 23 patients. With only 23 patients, the uncertainty around any performance estimate is large - the confidence intervals for all three models overlap substantially, meaning the apparent ranking could be reversed with more data.
The authors explicitly frame this as a Phase I exploratory study - the first stage in a multi-phase framework for evaluating medical AI. At this phase, the goal is simply to demonstrate feasibility: can the approach detect a signal? The answer appears to be yes, which justifies proceeding toward larger validation studies.
The fundamental limitation is sample size. With 112 total patients and a test set of only 23, the study is underpowered to demonstrate statistically conclusive differences between models or to generalize findings to broader populations. Many clinical AI studies face this challenge, particularly in specialized conditions where collecting large prospective datasets takes years.
An important validation step the researchers performed was a label shuffling test - they randomly reassigned responder/non-responder labels during training and re-trained the model. If the model performed similarly well with shuffled labels, it would suggest it was learning spurious patterns (called shortcut learning) rather than true biology. The shuffled labels model produced near-random performance (AUROC 0.49), confirming the original model's signal is genuinely related to the true labels.
One practical limitation is the reliance on expert pathologist annotation of AEH regions. In a future deployed system, requiring a pathologist to manually outline every biopsy slide before the AI can analyze it would make the workflow cumbersome. The authors plan to develop an automated region detection step to eliminate this manual prerequisite.
A reliable predictive test for hormonal treatment response would transform clinical decision-making for AEH and early endometrial cancer patients. For the 75% who would respond, it provides confidence that conservative hormonal management is the right choice, sparing them from surgery and allowing fertility preservation.
For the predicted non-responders, early identification enables a prompt pivot to more effective treatments - such as surgery or alternative medical options - before months are wasted on ineffective progestin therapy. This prevents both disease progression and the morbidity of prolonged hormonal side effects.
This kind of tool aligns with the broader movement toward precision medicine - matching treatments to the specific biological characteristics of each patient's disease rather than applying one-size-fits-all protocols. Just as genetic tests now guide targeted drug selection in many cancers, an AI-based pathology predictor could guide hormonal versus surgical treatment decisions.
The study also demonstrates how AI can analyze existing routine clinical materials - standard biopsy slides already collected at diagnosis - without requiring additional tests or procedures. The incremental cost of adding AI analysis to an already-collected biopsy could be modest compared to the potential benefit of personalized treatment allocation.
This FDA-led study represents the first published exploration of using deep learning on whole slide images to predict hormonal treatment response in AEH and early endometrial cancer patients - a clinically meaningful gap that has not previously been addressed with any validated biomarker.
The autoencoder model's AUROC of 0.80 demonstrates that meaningful biological signals predictive of treatment response are embedded in the pathological appearance of pre-treatment biopsy tissue. Future studies with larger, multi-institutional datasets are needed to validate and refine this signal.
The authors plan to incorporate multimodal data fusion - combining tissue image features with clinical factors like age, BMI, and race - which may further improve prediction accuracy. Modern AI architectures are well-suited to integrating these diverse data types into a unified prediction model.
Ultimately, if validated in larger prospective trials, a tool like this could become part of standard clinical workup at the time of AEH or low-grade endometrial cancer diagnosis, enabling personalized treatment planning from the very first appointment and setting a new standard for individualized cancer care.