Laboratory hematology sits at the center of diagnosing and monitoring the full spectrum of blood disorders, from acute leukemia and lymphoma to anemia, infections, and immunodeficiencies. The gold-standard diagnostic framework for hematologic malignancies is the MICM classification, which integrates results across four domains: morphology (M), immunophenotype (I), cytogenetics (C), and molecular biology (M). Together, these four streams inform diagnosis, risk stratification, treatment selection, and response monitoring. The problem is that nearly all of this work is still performed manually by human experts, making the system expertise-dependent, labor-intensive, and prone to interobserver variability.
The workforce gap: The MICM workflow demands hematopathologists and laboratory technicians with deep, accumulated clinical experience. This expertise is a scarce resource globally, and shortages are particularly acute in low- and middle-income settings. Even in well-resourced centers, manual interpretation creates long turnaround times (TAT) and limits the volume of specimens that can be processed. High workloads are directly associated with classification errors, particularly for uncommon or morphologically ambiguous cell types.
Escalating data complexity: The challenge has intensified as newer diagnostic technologies generate data of increasing dimensionality. Electronic health records now integrate laboratory results, imaging studies, and diagnostic codes. High-throughput genetic sequencing panels produce thousands of variant calls per patient. Spectral flow cytometry and mass cytometry (CyTOF) routinely measure 40 or more markers per cell in a single tube. Manual analysis of these high-dimensional datasets is not just inefficient but is approaching the limits of what unaided human cognition can reliably accomplish.
This 2025 review, authored by groups at West China Hospital (Sichuan University) and Tulane University, systematically surveys how AI, machine learning (ML), and deep learning (DL) are being applied across each sector of the MICM pipeline. The authors cover advances in CBC analysis, cytomorphology, flow cytometry immunophenotyping, cytogenetics, and molecular biology, while explicitly addressing limitations, deployment challenges, and future directions for each domain.
The complete blood count is the most commonly ordered laboratory test and often the first data point that raises clinical suspicion for hematologic malignancy. CBC outputs include counts of red blood cells (RBCs), white blood cells (WBCs), and platelets (PLTs), along with the WBC differential, volume distribution parameters, and automated flags. The International Society for Laboratory Hematology (ISLH) consensus guidelines (the "41 rules") recommend slide review whenever abnormal CBC parameters or flags are triggered, but this review step is labor-intensive and requires experienced hematologists. AI has been applied to make CBC-based malignancy detection faster, more sensitive, and less operator-dependent.
Acute leukemia subtype prediction: Haider et al. developed an artificial neural network (ANN) model using CBC and cell population data (CPD) parameters from 1,577 hematological neoplasm subjects to differentiate leukemia types, achieving 83.1% training accuracy and 89.4% test accuracy with AUC 0.79-0.94. Their subsequent work focused specifically on acute promyelocytic leukemia (APL), showing that platelet count, the immature platelet fraction (IPF), and neutrophil DNA/RNA content parameters could flag APL cases early. Alcazer et al. developed the AI-PAL XGBoost model, trained on CBC and biochemical parameters from six French university hospital databases, achieving validation set accuracies of 99.5% (ALL), 98.8% (AML), and 99.7% (APL) with AUC values of 0.97, 0.90, and 0.89 respectively. This model was externally validated and specifically designed to guide initial treatment decisions where cytological expertise is limited.
CML prediction from historical CBC data: Hauser et al. applied XGBoost and LASSO to historical CBC results from 1,623 patients, demonstrating that blood cell counts collected up to 5 years before a CML diagnosis could predict BCR-ABL1 positivity with AUC of 0.59-0.92, with basophils, leukocytes, and neutrophils identified as the most important predictors. A separate CNN-based approach using ResNet-18 mapped CBC scattergram images (rather than scalar parameters) to APL susceptibility, achieving precision above 0.99, sensitivity of 95%, and AUC above 0.99 with external validation, demonstrating that spatial patterns in automated analyzer output carry diagnostic information beyond the numeric readouts.
Prognosis and treatment monitoring: For diffuse large B-cell lymphoma (DLBCL), a retrospective multicenter LASSO model incorporating WBC count and hemoglobin alongside standard clinical variables outperformed conventional prognostic models (AUC 75.8% vs. 71.6% for random forest) for risk stratification. Across nine ML classifiers applied to 1,383 AML patients, hemoglobin at initial diagnosis predicted complete remission (AUC 0.77-0.86) and 2-year overall survival (AUC 0.63-0.75), with the ensemble of nine models proposed to improve methodological transferability across institutions.
Cytomorphology examination, the microscopic analysis of blood and bone marrow (BM) smears, is the most widely commercialized domain of AI-assisted hematology. The CellaVision digital microscope series, which uses ANN and DL algorithms, reduces total smear analysis time by approximately 50% and achieves 92% pre-classification accuracy compared to manual slide review. The DysplasiaNet CNN model, trained to detect dysplastic neutrophils (one of the most challenging morphological abnormalities), achieved sensitivity above 95%, specificity above 94%, and global accuracy of 94.85% when tested on 20,670 neutrophil images. The Sysmex DI-60, another ANN-based automated morphology analyzer, reduced hands-on time by 144.1 seconds per abnormal sample and 144.6 seconds per normal sample compared to manual review.
Peripheral blood (PB) classification benchmarks: The MC-100i CNN-based analyzer, evaluated in a multicenter study across 11 tertiary hospitals in China, achieved greater than 90% sensitivity and specificity for RBC classification and greater than 97% accuracy for normal WBC classification. Alferez et al. developed a system achieving 97.67% accuracy for classifying normal, reactive, and abnormal lymphoid cells across five groups. The MFDS-DETR model, combining multi-level feature fusion with deformable self-attention (a Transformer-based detection architecture), achieved AP 79.7% and AP50 97.2% for automated leukocyte localization and counting, with external validation in two public databases.
Bone marrow analysis: BM morphology automation is significantly harder than PB analysis because BM contains mixed cell classes across multiple maturation stages, and sample quality can vary substantially. A CNN trained on 171,374 expert-annotated single-cell images from 945 patients achieved automated classification of 22 classes of BM leukocytes with high precision and recall across cell types and external validation. A hierarchical patch-based DL framework using two Cascade R-CNN models achieved greater than 98.8% accuracy for BM whole-slide image analysis and identified megakaryocytes, mitotic cells, and four erythroblast stages in whole slide images with substantially reduced analysis time.
Disease-specific diagnostic systems: Kimura et al. developed the first automated MDS diagnostic system using PB smears, combining a CNN image recognition module with an XGBoost decision-making module. The system differentiated MDS from aplastic anemia with sensitivity 96.2%, specificity 100%, and AUC 0.990. The AMLnet DL pipeline discriminated not only between AML and healthy individuals but also between AML subtypes based on BM images at AUC above 0.95. A virtual hematological morphologist (VHM) framework incorporating three aspects of BM (cell conformation, hyperplasia degree, alkaline phosphatase score) achieved 99.23% balanced accuracy, 97.96% sensitivity, and 100% specificity for chronic-phase CML diagnosis, substantially outperforming end-to-end frameworks (97.11% vs. 68.75% generalization accuracy).
Multiparametric flow cytometry (FCM) is the primary method for immunophenotyping hematolymphoid malignancies and for measuring measurable/minimal residual disease (MRD). It identifies antigen expression on individual cells in a semiquantitative manner, making it indispensable for confirmatory diagnosis and treatment monitoring. The problem is that FCM data interpretation relies almost entirely on manual gating, a labor-intensive and subjective process in which operators draw sequential gates on bivariate scatter plots. With modern spectral FCM and mass cytometry (CyTOF) measuring 40 or more markers per cell simultaneously, manual gating of these high-dimensional datasets is approaching the limits of what is practical. The FlowCAP project (Flow Cytometry: Critical Assessment of Population Identification Methods), initiated by algorithm developers, FCM users, and instrument vendors, established benchmark datasets to compare automated clustering algorithms, concluding that automated methods are practical for many FCM use cases even if no single method is universally optimal.
Dimensionality reduction and visualization: The viSNE algorithm (visual interactive stochastic neighbor embedding) consistently discriminated BM samples from healthy individuals versus those with AML or ALL, and further distinguished newly diagnosed from relapsed leukemia samples. PhenoGraph defined phenotypes in high-dimensional single-cell data and revealed intracellular heterogeneity in pediatric AML, which was previously undetectable by standard manual gating. UMAP combined with random forest classification achieved greater than 95% accuracy across 3,417 PB cases for identifying key FCM parameters contributing to malignancy classification.
Leukemia diagnosis and subtyping: Monaghan et al. developed an ML model using data from 531 patients with cytopenia or acute leukemia that rapidly distinguished among APL, non-APL AML, ALL, and nonneoplastic cytopenia with 94.2% accuracy and 99.5% AUC. Notably, adding additional FCM parameters including light scatter properties and CD117 did not significantly improve the model, suggesting that a simplified antibody panel may suffice for malignancy screening. Zhao et al. classified nine diagnostic classes of mature B-cell neoplasms including CLL via a self-organizing map algorithm trained on 20,622 patients, with subsequent CNN classification achieving 83% accuracy for CLL. An RF classifier applied to 3,417 PB cases prospectively screened out normal cases and identified cases requiring add-on studies under a basal B-cell panel, achieving accuracy greater than 95% with 100% sensitivity.
MRD assessment: MRD positivity in acute leukemia and lymphoma is a powerful prognostic indicator, but detecting rare residual malignant cells among millions of normal cells is technically demanding. A hybrid deep neural network for CLL MRD measurement demonstrated 97.1% accuracy with excellent correlation to expert manual analysis. The DeepFlow algorithm, applied to 113 randomly selected CLL FCM files, showed highly consistent output with expert analysis while offering a simpler graphical interface and minimal batch effects compared to t-SNE or UMAP-based alternatives. DeepCyTOF automated cell type classification in CyTOF datasets with F-scores of 0.9921 and 0.9992 across two benchmark datasets without recalibration between datasets.
Cytogenetic analysis is vital in hematologic malignancy diagnosis because specific chromosomal abnormalities define disease entities and guide treatment. For example, the t(15;17)/PML-RARA fusion is diagnostic of APL and mandates ATRA-based therapy; t(8;21), inv(16), and t(16;16) define core-binding factor AML even when blast count is below 20%. The three conventional cytogenetic methods, karyotyping by chromosomal banding analysis, fluorescence in situ hybridization (FISH), and chromosomal microarray, each have inherent limitations: low resolution, limited analyzable metaphases, and inability to detect balanced rearrangements. Manual karyotype analysis involves photographing metaphase chromosomes, sorting them into 23 homologous pairs, and interpreting banding patterns, a process that is slow, expertise-dependent, and quality-sensitive.
Automated karyotyping: Haferlach et al. designed a fully automated deep neural network (DNN) classifier for chromosome assignment that achieved 98.6% accuracy across 23,000 chromosomes from patients with normal karyotypes. This system has been incorporated into ISO 15189-certified routine clinical workflow and is available cloud-based, representing one of the most advanced examples of AI integration into a clinical hematology laboratory setting. Vajen et al. subsequently developed a CNN predicting both chromosome class and orientation, achieving 98.8% accuracy for chromosomes from patients with AML, MDS, and chronic myeloproliferative disorder, with a 42% reduction in karyotyping workflow time. ChromoEnhancer (a CycleGAN model) transformed poor-quality karyograms into high-quality images with peak signal-to-noise ratio of 40.795 and structural index similarity of 0.988, without requiring a paired training set.
FISH and structural variant analysis: SpotLearn enabled fully automated, single-allele analysis of 3D genome organization in a high-throughput FISH workflow, and a subsequent DL pipeline localized and classified nuclei and FISH signals without segmentation at greater than 94% accuracy. For structural variations (SVs), the StrVCTVRE random forest classifier distinguished pathogenic from benign SVs overlapping exons with 83% accuracy and 90% sensitivity across multicenter evaluation. The Cue DL framework called and genotyped SVs with precision above 85% and recall above 87%, with external validation. For copy number variations (CNVs), the X-CNV XGBoost model achieved AUC 0.96 in training and 0.94 in validation for predicting CNV pathogenicity by integrating more than 30 predictive features genome-wide. A DL algorithm for chromothripsis detection in multiple myeloma patients achieved AUC 0.8309 with external validation, identifying a structural aberration associated with poor clinical outcomes.
Optical genome mapping (OGM): OGM is an emerging technology capable of imaging very long linear DNA molecules (median size above 250 kb) to detect SVs and CNVs with higher resolution than conventional karyotyping and shorter TAT. OGM has shown promise in AML, B-cell ALL, and CLL and has the potential to replace conventional karyotype, FISH, and chromosomal microarray with a single test. Current computational methods for OGM analysis fall short in accuracy and speed, presenting a clear opening for AI-based mapping algorithms.
Hematologic malignancies have historically been the vanguard in applying molecular genetics to cancer diagnosis and classification. The current WHO classification and the 2022 International Consensus Classification of Myeloid Neoplasms and Acute Leukemias rely heavily on genomic data, mandating identification of single nucleotide variants (SNVs), insertions and deletions (indels), structural variations, copy number variations, and gene fusions. Advanced sequencing modalities, including whole-exome, whole-genome, and RNA sequencing, generate datasets of enormous dimensionality and complexity that are increasingly beyond the reliable capacity of manual expert interpretation. AI is being applied at every step: variant calling, variant pathogenicity prediction, fusion gene identification, molecular subclassification, and prognostic model construction.
Variant calling and pathogenicity prediction: DeepVariant, a CNN-based tool that converts read pileup images around putative variants into genotype calls, achieved SNP F1 score of 99.95% and indel F1 score of 98.98% with external validation, and generalizes across diverse sequencing platforms (10x Genomics whole genomes, Ion Ampliseq exomes). An ML model trained on 1,372 SNVs and 939 indels in 87 genes predicted the clinical phenotype and outcomes of patients with panmyeloid leukemias (AML, MDS, chronic myelomonocytic leukemia, myeloproliferative neoplasms) with accuracy ranging from 63% to 88% across disease types. AlphaMissense, fine-tuned from AlphaFold, classified 89% of 71 million possible missense variants in the human proteome as likely benign or likely pathogenic, providing a comprehensive database of variant effect predictions available to clinical laboratories.
Gene fusion and molecular subclassification: An RF classifier combining middle-throughput gene expression data with ligation-dependent RT-PCR and NGS output (SNVs, gene fusions, and other markers) discriminated seven frequent categories of B-cell non-Hodgkin lymphomas with 80-100% concordance with prior classification results. Awada et al. integrated cytogenetic and gene sequencing data from 2,697 AML patients using an ML-driven genomic signature approach (Bayesian latent class analysis), achieving 97% accuracy for AML subtype classification and clinical outcome prediction, revealing novel genomic AML subclasses not previously recognized.
Prognostic model construction: Fleming et al. applied recursive partitioning and random forest to 2,074 non-APL AML patients, constructing a hierarchical prognostic risk model integrating cytogenetic and molecular factors that outperformed the European LeukemiaNet classification. Shreve et al. used genomic and clinical data from 3,421 AML patients to predict individualized outcomes, again demonstrating superiority over ELN classification. An ML model from 63 clinical and genomic variables from 2,043 MDS patients achieved concordance indices of 0.74 for overall survival and 0.81 for leukemic transformation, surpassing prior models by the same group. An interpretable ML model analyzing 24 commonly mutated MDS genes (including SF3B1, TET2, and ASXL1) distinguished MDS from other myeloid malignancies at diagnostic AUC of 0.951 with external validation and confidence intervals near 95%.
Data quality and retrospective design: The performance and robustness of AI in laboratory hematology is fundamentally constrained by data availability. Most published models are trained on retrospective, single-institution datasets. The authors invoke the maxim "garbage in, garbage out" to emphasize that even technically sophisticated models fail when trained on biased, non-representative, or poorly annotated data. Expert annotation for rare cell types (e.g., unusual blast morphologies, uncommon BM cell variants) requires scarce senior personnel time, creating annotation bottlenecks that limit training set diversity. Public, standardized cell image databases are insufficiently developed, and many institutions are reluctant to share patient data due to privacy concerns.
Lack of external and prospective validation: A recurring problem across all four MICM domains is the absence of rigorous external validation. Performance metrics from internal cross-validation routinely overestimate real-world accuracy. The few studies that have undergone prospective or multicenter validation often show meaningful performance drops, particularly in cytomorphology and FCM gating tasks where staining protocols, instrument settings, and panel designs vary between centers. For flow cytometry specifically, global variation in antibody panels, compensation settings, and instrument configurations means that models trained at one site cannot be assumed to generalize elsewhere without re-validation.
Panel and disease specificity in FCM: Most FCM AI models are designed for specific antibody panels and specific disease entities, limiting their applicability to the complex, overlapping disease presentations encountered in routine practice. A versatile model capable of simultaneous early screening, differential diagnosis, and MRD monitoring has not yet been developed. In cytogenetics, the lack of large public datasets with abnormal karyotype images (particularly for rare hematologic disorders) severely limits the development of AI classifiers for uncommon chromosomal rearrangements.
Integration and ethical barriers: Even well-validated tools face significant hurdles in clinical integration. These include the absence of standardized data formats across laboratory information systems, the technical complexity of deploying DL models in hospital infrastructure, and user interface limitations (many tools require programming expertise to operate). Ethical AI principles, including transparency, fairness, and privacy, must be embedded in algorithm development, particularly given that training datasets may over-represent certain populations and underperform for others. Regulatory approval pathways and liability frameworks for AI diagnostic tools remain underdeveloped in most jurisdictions.
Multimodal and comprehensive diagnostic frameworks: The authors argue that no single diagnostic modality provides a complete picture of hematologic malignancy, and the same limitation applies to AI models built on a single data type. The highest-value future direction is the development of AI systems that simultaneously integrate CBC parameters, morphological image data, FCM immunophenotyping results, cytogenetic findings, and molecular genetics into a single automated diagnostic overview. Such multimodal frameworks would mimic the comprehensive, multi-disciplinary tumor board review that currently defines best-practice hematologic diagnosis, but at the speed and scale that AI enables. Early conceptual work on "virtual hematological morphologists" combining multiple morphological and clinical inputs demonstrates both the promise and the complexity of this direction.
Point-of-care and resource-limited settings: AI integration into POC CBC devices represents a practical near-term opportunity to extend high-quality hematologic diagnosis to settings without access to large automated analyzers. The Sight OLO, validated in a multicenter study, requires only 27 uL of whole blood for a 19-parameter, 5-part differential CBC with AI-driven cell characterization, and supports fingerprick capillary blood collection. The Hilab system similarly combines ML and DL for CBC analysis from 90 uL whole or capillary blood, with correlation coefficients above 0.9 for most parameters compared to Sysmex XE-2100. These devices could democratize hematology screening in resource-limited clinical environments.
Prospective studies and interdisciplinary cooperation: Translation from proof-of-concept to clinical practice requires prospective studies that embed AI tools in active laboratory workflows and measure their impact on diagnostic accuracy, TAT, and patient outcomes. The authors emphasize that clinical trials and multicenter studies are prerequisites for AI models to be approved for routine use. These require extensive cooperation among computer scientists, clinicians, laboratory personnel, hospital administrators, and health policy makers. Frontline hematopathologists and laboratory technicians are specifically called upon not just to adopt AI tools but to actively help improve them by contributing domain knowledge, flagging edge cases, and validating model outputs against clinical findings.
Explainability and trust: Building clinical acceptance of AI tools requires not only high accuracy but interpretable outputs. For cytomorphology and flow cytometry, attention maps and SHAP values can indicate which image regions or parameter combinations drove a classification decision, allowing pathologists to interrogate the model's reasoning. In molecular biology, interpretable ML models that explicitly link genomic features to clinical outcomes provide actionable mechanistic insights rather than opaque predictions. The authors call for explainable AI to be designed in from the start rather than retrofitted, and for user-friendly graphical interfaces that do not require programming expertise as a prerequisite for clinical use.