Artificial intelligence's impact on breast cancer pathology: a literature review

Diagn Pathol 2024 Histopathology 8 Explanations View Original
Original Paper (PDF)

Unable to display PDF. Download it here or view on PMC.

Plain-English Explanations
Pages 1-2
AI in Breast Cancer Pathology: Why It Matters

Artificial intelligence (AI) has begun transforming clinical pathology by offering tools that can process large, complex histopathological images with speed and consistency that exceeds human capacity. In breast cancer, accurate pathological assessment underpins every treatment decision, from grading to receptor status to predicting chemotherapy response, making it a high-value target for AI assistance.

The demand for AI is partly driven by a growing gap between the volume of breast cancer cases and the available pathologist workforce. Breast cancer is the most common cancer among women in Europe and the United States, exceeding 10 percent of all cancers in several countries. As case numbers rise, tools that reduce repetitive manual tasks and improve diagnostic consistency become increasingly necessary.

AI in breast pathology falls into three main categories. Diagnostic AI identifies and classifies tumors and metastases from histological slides. Predictive AI quantifies biomarkers such as Ki-67, ER, PR, and HER2, and forecasts treatment response. Prognostic AI evaluates features such as tumor grade, mitotic count, and tumor-infiltrating lymphocytes to assess long-term patient outcomes.

This review from The Ohio State University synthesizes published research on open-source AI algorithms across all three categories, alongside a survey of commercially available AI platforms with regulatory approval in Europe or the United States, providing a comprehensive picture of the current state of AI in breast cancer histopathology.

TL;DR: AI tools for breast cancer pathology are organized into diagnostic, predictive, and prognostic categories and are motivated by growing case volumes, persistent inter-observer variability, and the potential to replace expensive molecular tests with computational image analysis.
Pages 3, 4, 5, 11
Diagnostic AI: Detecting Cancer, Invasion, and Metastasis

The HASHI framework (High-throughput Adaptive Sampling for Whole-slide Histopathology Image Analysis) was designed to detect invasive breast cancer in very large whole-slide images. Traditional CNNs struggle with the enormous number of pixels in WSIs, and HASHI addresses this by using adaptive sampling to train efficiently. Validated across nearly 500 cases and independently tested on 195 studies from The Cancer Genome Atlas, HASHI achieved a Dice coefficient of 76 percent, demonstrating robustness across different sites, scanners, and platforms.

For lymph node metastasis detection, the LYNA algorithm using Inception V3 architecture demonstrated significantly improved sensitivity compared to unassisted pathologists (91 percent versus 83 percent, p = 0.02) in a study of 130 cases reviewed by six pathologists. Algorithm-assisted review also reduced time per image and made classification easier for pathologists across experience levels.

The Visiopharm Integrator System metastasis AI algorithm was applied to screen lymph nodes in breast cancer patients and achieved 100 percent sensitivity and 100 percent negative predictive value, making it a promising triage tool for detecting metastatic deposits before detailed pathologist review. Its integration into clinical workflow reduced the need for immunohistochemical confirmation in some cases.

A study from Morocco by El Agouri et al. developed a deep learning approach using ResNet50 and Xception architectures on 328 digital slides, achieving 88 percent accuracy and 95 percent sensitivity in detecting breast carcinoma. The study demonstrated that effective diagnostic AI can be developed even with limited data in lower-resource settings, highlighting the global applicability of these tools.

TL;DR: AI systems have demonstrated strong performance in detecting invasive breast cancer and lymph node metastasis from whole-slide images, with several tools achieving sensitivity and accuracy that equals or surpasses unassisted pathologists.
Pages 5, 11
Predictive AI: Quantifying Biomarkers for Treatment Decisions

The Ki-67 proliferation index is a critical biomarker for breast cancer aggressiveness and treatment planning, but manual assessment by pathologists is highly variable. Abele et al. evaluated the Mindpeak Breast Ki-67 RoI tool across 204 slides reviewed by ten pathologists from eight institutions, achieving agreement rates of 95.8 percent for Ki-67 quantification and 93.2 percent for ER and PR. The tool demonstrated significant potential to reduce inter-observer variability across diverse clinical settings.

Bodén et al. compared visual estimation, fully automated digital image analysis (DIA), and a human-in-the-loop DIA approach for Ki-67 assessment. Visual estimation by pathologists performed significantly worse than DIA alone (p less than 0.05). Human-in-the-loop corrections primarily helped in cases with faint staining or poor tumor-stroma separation rather than uniformly improving results, suggesting that automated DIA is reliable for routine cases while human review adds value in edge cases.

For estrogen receptor (ER) scoring, the Visiopharm automated ER DIA algorithm was applied to 97 invasive breast carcinomas and achieved concordance of 93.8 percent with pathologist manual scoring. The fully automated digital workflow eliminated manual intervention and was shown to be feasible for direct clinical integration, saving time and labor while maintaining diagnostic reliability.

For HER2 scoring, the Visiopharm HER2-CONNECT application was used in 153 invasive breast carcinomas to correlate HER2 protein connectivity with response to anti-HER2 neoadjuvant chemotherapy. HER2 digital image analysis connectivity showed the strongest association with achieving pathological complete response, and in a validation cohort of 612 cases, concordance with pathologist manual scoring reached 87.3 percent with a 16 percent reduction in equivocal cases.

TL;DR: AI tools for Ki-67, ER, PR, and HER2 quantification consistently achieve better inter-observer agreement and higher concordance with expert pathologists than visual estimation alone, with fully automated workflows now feasible for clinical integration.
Pages 11-12
AI for Predicting Neoadjuvant Chemotherapy Response from Biopsies

Shen et al. developed a multi-pipeline AI system using CNN analysis with the ResNeXt model, support vector machines, and random forest classifiers applied to H&E images of pre-chemotherapy needle biopsies. The system analyzed three independent models focusing on different cancer atypia characteristics and achieved 95.15 percent accuracy in predicting neoadjuvant chemotherapy response in a test set of 103 cases, demonstrating that nuclear phenotypes visible in pre-treatment biopsies encode response information.

Aswolinskiy et al. introduced interpretable computational biomarkers from H&E slides called PROACTING, derived from tumor-infiltrating lymphocyte assessment, segmentation-based features, and mitotic count. Using a training set of 721 patients and validation set of 126, the approach achieved AUC values between 0.66 and 0.88 for predicting pCR. The TILs biomarker showed particularly promising sensitivity in TNBC and Luminal B cohorts, and the authors emphasized interpretability as a key advantage over black-box deep learning models.

Saednia et al. built a hierarchical deep learning framework combining CoAtNet at the patch level with Vision Transformer (ViT) models at the tumor level and a patient-level prediction module. Trained on 144 patients with 9,430 annotated tumor regions and validated on 63 patients, the hierarchical model outperformed two-level and patch-level-only approaches, achieving AUCs of 0.79, 0.81, and 0.84 with F1-scores of 86, 87, and 89 percent across three analysis settings.

The IMPRESS pipeline by Huang et al., which quantifies spatial immune microenvironment features from paired H&E and IHC whole-slide images, achieved AUC of 0.90 in HER2-positive breast cancer and 0.77 in TNBC for predicting pCR. Unlike pure deep learning approaches, IMPRESS features are directly tied to specific immune markers in specific tissue regions, enabling interpretable output that pathologists can verify and understand. Whitney et al. separately showed that computerized nuclear morphology analysis could predict Oncotype DX risk categories with 75 to 86 percent accuracy, suggesting AI may partially replace expensive molecular tests.

TL;DR: Multiple AI approaches applied to pre-treatment biopsy H&E and IHC images can predict neoadjuvant chemotherapy response with accuracy ranging from 79 to 95 percent, with interpretable biomarker-based systems offering advantages over pure black-box deep learning approaches.
Pages 6, 12, 13
Prognostic AI: Grading, Mitosis Counting, and TIL Assessment

Nottingham histological grading is a cornerstone of breast cancer prognosis but is subject to inter-observer variability, particularly for grade 2 tumors that share features with both low- and high-grade cancers. Wang et al. developed DeepGrade, a deep CNN model trained on 1,567 cases and tested on 1,262, specifically to reclassify grade 2 (NHG2) tumors into lower- and higher-risk subgroups. DeepGrade-high NHG2 patients showed significantly increased risk for recurrence (HR = 1.91, p = 0.019), and the method predicted prognosis as effectively as gene expression profiling while using only routine stained tissue sections.

Mitotic figure counting is a required component of breast cancer grading but is time-consuming and variable. Pantanowitz et al. evaluated AI-assisted mitosis counting using a ResNet-101 RCNN in 320 invasive ductal carcinoma cases across readers of varying expertise. AI assistance improved accuracy from 43.9 percent to 55.2 percent, reduced false positives, and saved 27.8 percent of counting time. Critically, benefits were observed across all experience levels, suggesting AI assistance is valuable for pathologists regardless of seniority.

Tumor-infiltrating lymphocytes (TILs) are an established prognostic biomarker in triple-negative breast cancer, with higher TIL density associated with better outcomes. Balkenhol et al. used automated deep learning to assess CD3, CD8, and FOXP3 markers across different tumor regions in 94 TNBC specimens. TIL abundance showed consistent negative correlations with recurrence-free and overall survival across all markers and measurement regions, confirming TILs as robust prognostic biomarkers. Importantly, the study found that limiting TIL assessment to tumor stroma alone, as per international guidelines, did not compromise prognostic value.

For phyllodes tumors, accurate mitosis counting is important for grading but challenging due to tumor rarity. Chow et al. compared counting in 10 high-power fields versus whole-slide image methods for 93 cases and found both correlated equally with tumor grade and stromal features. Neither method correlated with patient age or tumor size. The study highlights the need for standardized methods for mitosis counting on whole-slide images across tumor types.

TL;DR: AI improves consistency in breast cancer grading, mitosis counting, and TIL assessment, with DeepGrade identifying high-risk grade 2 tumors invisible to conventional methods and AI-assisted mitosis counting saving nearly 30 percent of counting time while improving accuracy.
Pages 7, 8, 9, 10, 15, 16
Commercial AI Platforms for Breast Pathology

Mindpeak (Hamburg, Germany) offers a CE-IVD certified suite of tools including Mindpeak Breast HER2 RoI, Breast Ki-67 HS, and Breast ER/PR for automated quantification of standard breast cancer biomarkers from immunohistochemistry slides. These tools operate on whole-slide images and require an interactive image viewer, producing standardized scores that reduce subjective variability in biomarker reporting.

Visiopharm offers a comprehensive portfolio including CE-IVD certified tools for ER, PR, Ki-67, HER2-CONNECT, invasive tumor detection, lymph node metastasis detection, HER2-SISH, and HER2-FISH analysis. The Visiopharm HER2-CONNECT tool specifically predicts response to anti-HER2 neoadjuvant chemotherapy. The lymph node metastasis detection app demonstrated 100 percent sensitivity in published validation, making it a potentially transformative screening tool for pathology workflows.

PathAI's AIM-HER2 tool produces both a HER2 score and a density heatmap using additive multiple instance learning (aMIL), making its reasoning transparent to pathologists. It is compatible with multiple scanner vendors including Leica Aperio and Hamamatsu NanoZoomer. Paige offers the CE-IVD and UKCA certified Breast Suite and Paige Breast Lymph Node tools for H&E-based detection of premalignant and malignant neoplasms and lymph node metastasis, respectively, both operating on Leica Aperio scanners with a clinical-grade viewer.

IBEX GALEN Breast provides CE-IVD certified AI-powered pathology for breast cancer subtype detection and grading from H&E slides and is compatible with any whole-slide image scanner. OWKIN RlapsRisk BC, distributed through Mindpeak, uses H&E slides to predict the risk of distant relapse in ER-positive, HER2-negative invasive carcinoma patients, representing a new category of AI tools that generate prognostic risk scores directly from routine stained slides without requiring molecular testing.

TL;DR: Multiple commercially available AI platforms with regulatory certification now cover the full range of breast pathology tasks from biomarker scoring to metastasis detection to chemotherapy response prediction, with deployment requiring whole-slide imaging infrastructure and compatible viewers.
Pages 13-15
Limitations of Current AI Tools in Breast Pathology

The most consistent limitation across all published AI tools is sensitivity to preanalytical variables. Poor sample quality, staining artifacts, air bubbles, and unfamiliar staining patterns caused both pathologists and AI algorithms to fail in equivalent ways. Since AI models learn patterns from training data, any systematic difference between training and deployment conditions, whether due to staining protocol, scanner type, or tissue processing, can degrade performance unpredictably.

Manual annotation requirements remain a bottleneck. Detection of lymph node metastasis requires marking extensive pixel-level regions in each image, which is labor-intensive and limits scalability. NAC prediction pipelines often require manual sampling of regions of interest from annotated cancer areas. Automated annotation tools for complex lesions such as ductal carcinoma in situ (DCIS) are not yet available, requiring expert manual labeling that constrains the size of training datasets that can realistically be assembled.

Standardization gaps affect mitosis counting specifically. There is no established agreement on what area of a whole-slide image should be used for mitosis counting, with different studies using varying sizes. Without consistent protocols, AI models trained in one setting cannot be directly compared to or deployed in another. The recommended standard of counting multiple screens at 40x magnification for a 3 mm2 area equivalent to 10 high-power fields is currently a practical guideline rather than a validated international standard.

For TIL assessment, a lack of uniform methods across studies limits comparability and generalizability. The prognostic studies conducted to date focused predominantly on TNBC, and whether findings translate to other breast cancer subtypes is unknown. Automated segmentation of tumor cells for TIL analysis was not optimal for all tumor morphologies, introducing noise into measurements. These gaps underscore the need for larger, subtype-stratified prospective validation studies before computational TIL assessment can be standardized for clinical use.

TL;DR: Current AI tools in breast pathology are limited by sensitivity to preanalytical variables, heavy annotation requirements for model training, lack of standardized protocols for mitosis counting, and insufficient validation across breast cancer subtypes beyond TNBC.
Pages 14-15
Future Directions and Summary

The most impactful near-term opportunity is developing AI systems capable of automatically identifying nuclear phenotypes associated with chemotherapy response without requiring manual sampling of regions of interest. Full automation of pre-treatment biopsy analysis pipelines would enable scalable deployment of NAC prediction tools in routine clinical practice, supporting personalized treatment decisions without added pathologist workload.

For prognostic applications, larger validation studies are needed to determine which TIL assessment approach, which markers, and which measurement regions provide the most reliable prognostic information across breast cancer subtypes. Integration of multiplex immunohistochemistry with deep learning-based spatial analysis, as demonstrated by the IMPRESS framework, represents a promising path toward richer and more reproducible prognostic profiling.

AI integration with genomic data such as Oncotype DX risk scores offers a compelling direction for improving risk stratification without expensive molecular tests. If AI-derived morphological features can reliably correlate with intermediate-risk Oncotype DX categories, it may be possible to spare some patients from costly genetic testing while achieving equivalent treatment guidance, representing both a cost-saving and an access-improving development.

This review concludes that AI has already demonstrated meaningful contributions to breast cancer pathology across diagnostic, predictive, and prognostic domains. Continued research addressing preanalytical variability, annotation scalability, and multicenter prospective validation is essential to translate promising research findings into robust, standardized clinical tools that can reduce avoidable diagnostic errors and support personalized treatment decisions at scale.

TL;DR: Advances in AI-driven breast pathology hold genuine transformative potential, with the most immediate priorities being full automation of NAC prediction pipelines, standardized TIL assessment protocols, and validation of AI as a cost-effective alternative to molecular risk stratification tests.
Citation: Open Access, 2024. Available at: PMC10882736.