Lymphomas rank among the ten most common cancers globally and present clinicians with one of the most diagnostically demanding environments in all of oncology. The current WHO classification recognizes approximately 100 distinct subtypes of lymphoid neoplasms, each requiring accurate identification for appropriate treatment selection. Despite years of training, even experienced pathologists face real-world error rates that carry direct patient consequences. A landmark French nationwide study through the Lymphopath Network found that roughly 20% of lymphoma diagnoses issued by nonexpert pathologists were inaccurate, a figure that translated directly into inappropriate treatment decisions.
The data generation problem: Modern haematopathology is being reshaped by two convergent trends: the proliferation of digital whole-slide imaging (WSI) systems and the expansion of high-throughput molecular sequencing. Both are generating data volumes far exceeding what pathologists can interpret through visual inspection alone. Fluorescence in situ hybridization (FISH), clonality assays, RNA sequencing, methylation profiling, and whole-exome or whole-genome sequencing are now routinely applied to lymphoma specimens, but the informatics infrastructure to synthesize results across these modalities remains immature.
Scope and objectives of the review: Published in Histopathology in 2025 (with 2024 open-access publication), this review by Syrykh, van den Brand, Kather, and Laurent provides a structured overview of AI applications in haematopathology. The authors organize the field into two major categories: diagnostic assistance tools (H&E slide analysis, immunohistochemistry and FISH image quantification, genomic data interpretation) and exploratory approaches (predicting molecular alterations from histology, prognosis prediction, treatment response forecasting, and multimodal integrative models). The review draws on published studies across multiple institutions and countries, reflecting the international state of the science rather than any single center's experience.
A key framing argument is that AI in haematopathology should not aim to replace the pathologist but to complement expertise, reduce interobserver variability, and automate time-consuming quantification tasks. The review also acknowledges the significant gap between research performance and clinical deployment readiness, noting that at the time of writing, only approximately 50 AI applications held CE certification in Europe and just six pathology applications had received FDA approval in the United States, with very few dedicated to haematopathology.
The review opens its technical section with a clear taxonomy of AI approaches relevant to pathology. Machine learning (ML) encompasses algorithms that learn patterns from data through iterative exposure rather than explicit programming rules. Three development stages apply to any ML model: selecting and preparing training data, choosing an algorithm appropriate for the data type and problem structure, and iteratively training the algorithm by comparing its predictions against known outputs to adjust weights and improve accuracy. The standard split of datasets into training, validation, and test subsets is critical for avoiding overfitting and obtaining unbiased estimates of real-world performance.
Three types of machine learning: Supervised learning uses labeled data where each training example has a known output (e.g., a slide labeled as follicular lymphoma vs. DLBCL). Unsupervised learning discovers structure in unlabeled data, which requires large datasets but can reveal novel clusters or subtypes without human annotation. Reinforcement learning trains agents to maximize rewards through trial and error; it is less commonly applied in pathology but is relevant for sequential decision problems such as treatment sequencing. The spectrum from fully human-guided to fully machine-guided approaches maps across these three categories, and the choice of algorithm (regression, decision tree, random forest, support-vector network, neural network) depends on the data volume and problem complexity.
Deep learning as a subset of ML: Deep learning uses multilayer neural networks where each layer transforms the input representation into progressively more abstract features. The input layer receives raw data (image pixels, genomic reads, or clinical variables); hidden layers extract increasingly high-level patterns; and the output layer assigns probabilities to each class. Convolutional neural networks (CNNs) are particularly well-suited to image analysis because convolutional filters can detect local spatial patterns (cell morphology, tissue architecture) at multiple scales simultaneously. The depth of these networks is what enables them to detect subtle features that are undetectable to the human eye, such as architectural patterns correlating with aggressive transformation in CLL.
A recurring theme throughout the review is that the performance of any ML model is only as good as the training data it learns from. Representative, diverse, and carefully curated datasets are essential to avoid bias, and training data from a single institution or scanner type may not generalize to different clinical settings. This foundational constraint shapes every application discussed in subsequent sections.
The most mature AI application in haematopathology involves automated classification of digitized H&E-stained WSIs. Published deep learning models for lymphoma detection and subtype classification consistently report area under the receiver operating characteristic curve (AUROC) values exceeding 0.90, with some studies on binary classification tasks (e.g., DLBCL vs. non-DLBCL, or lymphoma vs. reactive tissue) reaching AUROC above 0.95. These algorithms learn directly from high-resolution tissue images, identifying morphological features such as cell size, nuclear irregularity, and architectural growth patterns that correlate with specific diagnostic categories.
Current coverage and its limits: A critical limitation acknowledged by the review is that published AI models overwhelmingly focus on the three most common lymphoma subtypes: DLBCL, follicular lymphoma (FL), and chronic lymphocytic leukemia/lymphocytic lymphoma (CLL). None of the currently published algorithms approach the diagnostic breadth of a practicing haematopathologist, who must be capable of recognizing nearly 100 different lymphoid subtypes plus a wide variety of reactive lesions in daily practice. The rarity of many subtypes creates a fundamental data scarcity problem: building the large, annotated training datasets needed for rare entities requires multicenter collaboration at a scale that few research groups have achieved.
Beyond binary classification: More nuanced AI capabilities are also being demonstrated. AI tools have been applied to grade assignment in follicular lymphoma, where interobserver variability among pathologists is well-documented and directly affects treatment decisions. Models targeting the detection of accelerated phase or transformation of CLL (a clinically critical event with significant prognostic implications) have shown that AI can extract morphological and architectural biomarkers that are difficult for humans to quantify consistently. The review argues that AI should not merely replicate existing diagnostic criteria but also discover new morphological biomarkers predictive of clinical behavior that have not previously been identified.
Practical deployment requires that AI predictions be accompanied by a confidence index. Pathologists should not be expected to act on low-confidence predictions; the confidence metric allows the algorithm to flag uncertain cases for expert review rather than forcing a categorical output. This design principle is essential for safe integration of AI into diagnostic workflows.
Lymphoma diagnosis depends not only on H&E morphology but on an extensive panel of immunohistochemical (IHC) markers and, for specific entities, on FISH analysis to detect chromosomal rearrangements. These ancillary techniques are labor-intensive and subject to interpretation variability between pathologists and between laboratories. Automated image analysis tools offer a route to standardizing quantification, reducing subjectivity, and potentially identifying new biomarkers that conventional assessment cannot reliably detect.
Ki67 as a case study: The Ki67 proliferative index is one of the most clinically important prognostic markers in mantle cell lymphoma (MCL). It is incorporated directly into the biologic-Mantle Cell Lymphoma International Prognostic Index (MIPI-b) score, which guides treatment decisions. However, the methodology for quantifying Ki67 staining and the cutoff values used to stratify patients vary substantially between studies and institutions, at least in part because manual assessment is inherently subjective. AI-based image analysis can standardize Ki67 quantification by applying a consistent computational method across all cases, and machine learning approaches can then optimize the threshold values for patient stratification on large datasets rather than relying on expert consensus from limited samples.
Cell-of-origin classification in DLBCL: Beyond single-marker quantification, ML-based algorithms that integrate multiple IHC markers simultaneously have outperformed the gold-standard Hans algorithm for classifying DLBCL into germinal center B-cell (GCB) and non-GCB subtypes (cell-of-origin, or COO classification). Several independent studies have demonstrated that ML approaches analyzing IHC marker combinations achieve better performance compared to the rule-based Hans algorithm, which uses a binary decision tree with fixed cutoffs. Because COO classification has direct implications for prognosis and trial eligibility under some treatment protocols, this improvement has practical clinical significance.
Digital FISH analysis: The digitization of FISH slides has enabled automated capture and analysis systems that save interpretation time, standardize counting methods and positivity thresholds, and allow analysis of larger numbers of nuclei per case than manual review permits. Beyond counting efficiency, automated FISH tools can compensate for truncated nuclei in tissue sections, detect aneuploidies, and identify alternative fusion partners that might be missed in manual review. These capabilities are particularly important for FISH analyses used to identify chromosomal rearrangements diagnostic of Burkitt lymphoma and high-grade B-cell lymphoma with MYC/BCL2/BCL6 rearrangements.
High-throughput sequencing has become central to lymphoma classification, generating large numbers of variants whose clinical significance must be manually assessed by pathologists and molecular biologists. Variants are categorized on a five-tier scale (pathogenic, probably pathogenic, variant of uncertain significance, probably benign, benign), and the classification process integrates epidemiological databases, clinical criteria, co-occurrence data, predictive bioinformatics, and functional studies. This process is time-consuming and prone to interobserver variability, particularly for variants of uncertain significance, which represent a large fraction of results in lymphoma sequencing panels.
The lymphoma-specific challenge: Unlike solid cancers, where genomic variants often serve as direct therapeutic targets (e.g., EGFR mutation in lung cancer), molecular data in lymphomas are used primarily for classification rather than targeted therapy selection, and no practice guidelines exist for interpreting the clinical relevance of genomic variants in the lymphoma context. This absence of consensus, combined with the scarcity of large reference databases for rare lymphoma entities, makes the interpretation problem especially prone to inter-laboratory variability. AI approaches that learn from large collections of annotated variant calls could standardize interpretation and reduce the burden on individual pathologists.
Deep learning for variant calling: Several deep-learning-based variant calling tools have been developed to improve somatic mutation detection from sequencing data, outperforming conventional bioinformatics pipelines in accuracy for complex variants. In the context of lymphoma subtype classification from gene expression data, Bobee et al. trained a random forest algorithm to discriminate the seven most frequent B-cell NHL entities using transcriptomic profiles, demonstrating that gene expression data alone can distinguish lymphoma subtypes with clinically useful accuracy. Zhang et al. extended this approach with an AI-driven bioinformatics system identifying non-mutually exclusive genetic signatures in DLBCL, going beyond simple binary subtyping toward a more continuous molecular landscape.
The main practical limitations are data scarcity and the limited adoption of high-throughput sequencing as standard of care. Most studies focus on DLBCL because it offers the largest available molecular databases. RNA sequencing, methylation analysis, and copy number variation analysis all represent additional layers of molecular data that AI can integrate, but their clinical use remains heterogeneous across centers. Expansion of genomic databases and wider adoption of sequencing in routine practice are prerequisites for realizing the full potential of AI in lymphoma molecular diagnostics.
One of the most scientifically ambitious applications of AI in haematopathology is using H&E slide morphology to predict underlying genetic alterations, bypassing the need for separate molecular testing. The conceptual basis is that morphology reflects the genetic and epigenetic makeup of tumor cells: certain genetic alterations cause recognizable changes in cell shape, nuclear size, chromatin texture, and tissue architecture that, while subtle, can be learned by deep learning models. If accurate enough, such models could function as low-cost screening tools that flag cases for confirmatory molecular testing only when a genetic alteration is likely present.
Performance in solid tumors and lymphomas: In solid tumors, proof-of-concept results have been impressive. In non-small cell lung cancer, deep learning models predicted the presence of specific common mutations from H&E slides with AUROC values of 0.73-0.86. In gastrointestinal carcinoma, models predicting microsatellite instability (MSI) from histology achieved AUROC up to 0.96 in large international cohorts, and CE-approved MSI detection algorithms are now commercially available for colorectal carcinoma. For lymphoma specifically, three independent studies in DLBCL applied machine learning to H&E morphology to predict MYC rearrangement, yielding AUROC values of 0.68-0.83. A separate single-center study in 57 large B-cell lymphoma cases predicted double/triple hit rearrangements (concurrent MYC plus BCL2 and/or BCL6) with sensitivity of 100% and specificity of 87%, though the small sample size limits the generalizability of this result.
Prognosis prediction from gene expression: AI models trained on gene expression profiles have demonstrated consistent ability to predict outcomes in lymphoma. In 184 follicular lymphoma patients, AI analysis identified a 43-gene expression signature predictive of overall survival. For DLBCL, AI has been applied to predict clinical course, response to bortezomib-based regimens, and cell-of-origin subtype. A model combining morphology, immunophenotype, and clinical features in high-grade B-cell lymphoma (Kong et al.) identified independent risk factors and showed that first-line R-CHOP had poor efficacy in this aggressive subtype. Another study using national lymphoma registry data demonstrated better prognostic performance than the International Prognostic Index (IPI) score when an ML model was trained on a large nationwide cohort.
These exploratory results collectively illustrate that AI can extract prognostically relevant information from data sources, both morphological and molecular, that currently contribute nothing to formal risk stratification. The translation pathway from research finding to clinical tool requires prospective validation in diverse cohorts, but the conceptual framework is well-established.
Data scarcity and subtype imbalance: The rarity of individual lymphoma subtypes is the most pervasive technical obstacle. Neural networks require large, balanced training datasets to generalize reliably, but some lymphoma entities occur in fewer than one in a million individuals, making it impossible for any single institution to assemble a training cohort of meaningful size. Even for common entities like DLBCL, the heterogeneity of the disease means that models trained at one center may not generalize to the molecular or morphological spectrum represented at another. Federated learning, where model weights are shared across institutions without pooling patient data, is presented as a promising solution but remains under active development for haematopathology applications.
Digitization infrastructure: Clinical deployment of AI tools for WSI analysis requires that laboratories have already transitioned to digital pathology workflows, including whole-slide scanners, image management systems, and the computational infrastructure to run inference models on large images. Many pathology laboratories globally, particularly in lower-resource settings, have not completed this transition, creating an infrastructure gap between research possibility and clinical reality. The review estimates that clinical implementation of AI in digital pathology will accelerate in coming years but acknowledges the current lag.
Explainability and the black-box problem: Many deep learning models, particularly those operating directly on image data, cannot explain their predictions in terms that pathologists can evaluate or that align with established diagnostic criteria. While attention map visualization techniques (such as Grad-CAM) can highlight which regions of a slide contributed to a prediction, these heatmaps do not constitute a mechanistic explanation equivalent to what a pathologist would provide. The lack of interpretability undermines clinical trust and creates liability concerns when AI outputs are used in patient care decisions. The review notes that explainability methods are being increasingly incorporated into AI tools but that the problem remains fundamentally unsolved.
Regulatory and ethical gaps: Deploying AI in clinical decision-making raises questions about liability when AI outputs are incorrect, the appropriate balance between automated decision support and physician oversight, and the privacy implications of training models on patient data. The absence of standardized reporting guidelines for AI studies in haematopathology makes comparative evaluation of published models difficult, and the regulatory pathway from research tool to approved clinical software is slow and resource-intensive. The review identifies regulatory framework development as an essential prerequisite for safe clinical deployment.
Foundation models and self-supervised learning: The review highlights foundation models as the most transformative near-term development for computational haematopathology. Foundation models are large, task-agnostic AI models trained on massive unannotated datasets using self-supervised learning objectives. They learn generalizable feature representations of tissue morphology that can be adapted to specific downstream tasks (lymphoma subtype classification, mutation prediction, prognosis estimation) using only small amounts of task-specific labeled data. Crucially, their robustness across different staining protocols, scanner hardware, and patient demographics makes them far better suited to real-world deployment than models trained from scratch on single-center datasets. Several whole-slide image foundation models have been published, including those trained on millions of pathology tiles from large cancer center repositories.
Generative AI for data augmentation: Generative AI approaches can create synthetic histopathology images that realistically simulate rare lymphoma subtypes, potentially alleviating the data scarcity problem for training supervised classifiers. By generating plausible variations in staining intensity, tissue processing artifacts, and morphological features, generative models can expand training datasets for rare entities without requiring additional biopsy material. The review notes that this application holds significant promise but that clinical integration is still hampered by limitations in data availability, computational demands, interpretability, and regulatory frameworks.
Natural language processing and clinical data integration: NLP tools using text mining algorithms can extract and structure information from free-text pathology reports, enabling the construction of large, well-annotated retrospective datasets from existing laboratory information systems. In haematopathology, where diagnoses, molecular results, and clinical outcomes are often recorded in narrative form across multiple report types, NLP offers a practical route to converting unstructured data into training inputs for predictive models. NLP can also assist pathologists in generating accurate diagnostic reports and improve interoperability between information systems across institutions.
Toward clinical integration: The review outlines a staged implementation model for AI in haematopathology. Initial deployment should target high-volume, routine diagnostic tasks (such as flagging cases likely to require expert second-opinion review) before expanding to complex classification and prediction tasks. Pilot programs in digitized laboratories should be followed by prospective validation studies measuring AI impact on diagnostic accuracy, time-to-diagnosis, and patient outcomes. The pathologist role in this future model is to validate algorithm performance and control outputs rather than to be replaced by them. The authors emphasize that the most promising AI applications will be those designed in explicit partnership with practicing haematopathologists.