Radiology sits at the intersection of digital medicine and artificial intelligence: the combination of vast archived imaging datasets and affordable high-performance computing has made it the medical specialty most rapidly transformed by machine learning. Within hematology, ML has been applied in adjacent areas including cytomorphometry (characterizing cell populations by morphology), cytogenetics (identifying chromosomal abnormalities), and immunophenotyping (flow cytometry cell classification). But a comprehensive assessment of ML applied specifically to radiological imaging of hematological malignancies had not previously been performed, which is the gap this scoping review from Kotsyfakis et al. (published in Frontiers in Oncology, December 2022) sets out to fill.
Clinical imperatives for diagnosis: Many hematological malignancies present imaging challenges that create genuine diagnostic dilemmas. The most prominent is distinguishing primary central nervous system lymphoma (PCNSL) from glioblastoma multiforme (GBM) on MRI. While GBM typically shows ring-like or heterogeneous enhancement with hypointense necrotic cores, and PCNSL typically shows uniform enhancement with low cerebral blood volumes (CBVs), atypical presentations overlap significantly. Non-necrotic GBMs and necrotic PCNSLs can look indistinguishable, and "hypervascular PCNSLs" with elevated CBVs mimic GBM. This distinction is clinically critical because PCNSL is treated with methotrexate-based chemotherapy while GBM requires surgical resection followed by radiochemotherapy. A biopsy can resolve the dilemma but is invasive and may be non-diagnostic if steroids have already lysed lymphoma cells.
Clinical imperatives for segmentation: Total metabolic tumor volume (TMTV), the quantification of metabolically active tumor assessed by FDG-PET/CT, is an established prognostic biomarker for both Hodgkin and non-Hodgkin lymphomas. However, computing TMTV currently requires manually marking many regions of interest (ROIs), a time-consuming and operator-dependent process that introduces error. Multiple non-standardized thresholding approaches are in use (SUV 41%, SUV 2.5, or SUV mean liver uptake), which impedes cross-study comparability. Automated, standardized ML segmentation would be valuable for both clinical efficiency and research reproducibility.
Clinical imperatives for prognostication: Only approximately 60% of diffuse large B-cell lymphoma (DLBCL) patients benefit from current first-line therapies, and roughly 15% experience primary treatment failure with a median survival under one year after that point. Identifying upfront who will fail R-CHOP would enable tailored intensification with emerging therapies such as CAR-T cell therapy. Beyond tumor bulk, intratumoral heterogeneity reflects molecular and microenvironmental differences that drive progression and therapeutic responses, and these spatial patterns in imaging data are precisely the kind of high-dimensional information that ML is well-positioned to extract.
The authors followed the Preferred Reporting Items for Systematic Reviews and Meta-Analysis Extension for Scoping Reviews (PRISMA-ScR) guidelines. A scoping review methodology was chosen over a systematic review with meta-analysis for two reasons: the available evidence was still emerging, making a first-impression mapping appropriate, and the included studies were highly heterogeneous in their ML methods, imaging modalities, disease entities, and outcome measures, which precludes statistical pooling. The scoping approach allowed the authors to map available evidence, clarify key concepts, characterize how research was being conducted, and identify knowledge gaps.
Search strategy: The PubMed database was searched from inception to October 1, 2021. The search term combined ML-related terms (machine learning, artificial intelligence, decision tree, neural network, random forest, support vector machine, radiomics) with imaging terms (radiology, imaging, tomography, magnetic resonance) and hematological malignancy terms (hematological malignancy, lymphoma, myeloma, leukemia). The inclusion criteria followed a population-concept-context framework: (i) pediatric and adult patients with suspected or confirmed hematological malignancy undergoing imaging; (ii) any study using ML techniques to derive models from radiological images for clinical benefit; and (iii) original research articles from any global setting. Commentaries, editorials, letters, and case reports were excluded. Importantly, all modeling approaches defined as ML in the respective papers, including logistic regression, were included as long as the analysis was substantially computer-driven.
Quality assessment tools: Because no quality assessment tool at the time of the review specifically addressed ML methodology, the authors applied two separate frameworks. For diagnostic and segmentation studies, Quality Assessment of Diagnostic Accuracy Studies 2 (QUADAS-2) criteria were operationalized using relevant Checklist for Artificial Intelligence in Medical Imaging (CLAIM) items, assessing four domains: patient selection risk of bias, index test risk of bias, reference standard risk of bias, and applicability concerns. For prognostic and predictive studies, the Newcastle-Ottawa Scale (NOS) was used, with scores converted to Agency for Healthcare Research and Quality (AHRQ) standards categorizing studies as good, fair, or poor quality. The review acknowledges that no current tool fully addresses the specific methodological demands of ML model development.
Data extraction: From papers meeting inclusion criteria, the authors extracted study population characteristics, imaging modalities, ML algorithms used, model validation methods, performance measures (accuracy, sensitivity, specificity, AUC), and direct comparisons with radiologist performance or other algorithms. Of 397 initially identified studies, 53 met inclusion criteria. The most common exclusion reasons were that ML was not the primary analytical methodology or that studies used features from non-imaging data such as histopathological or cytology images.
Thirty-three of the 53 included studies applied ML to diagnose a hematological malignancy or distinguish it from another disease entity. The dominant application, accounting for 18 of these 33 studies, was discriminating gliomas (predominantly GBM) from PCNSL using features extracted from MRI (17 studies) or FDG-PET (one study). The remaining 15 studies divided between those differentiating other solid hematological malignancies from benign or malignant lesions at other sites (nasopharyngeal carcinoma from nasopharyngeal lymphoma, idiopathic orbital inflammation from ocular adnexal lymphoma, thymic neoplasms from thymic lymphoma, breast carcinoma from breast lymphoma, lymphoma from normal lymph nodes, multiple myeloma from bone metastases) and those detecting the location or presence of hematological malignancies (lymphoma lesion localization, leukemia bone marrow involvement, multiple myeloma bone marrow infiltration, mantle cell lymphoma).
Algorithm diversity: A wide array of ML approaches was deployed across these studies. For GBM versus PCNSL discrimination, algorithms included support vector machines (SVMs), linear discriminant analysis (LDA), logistic regression (LR), artificial and convolutional neural networks (A/CNNs), k-nearest neighbors (KNNs), naive Bayes classifiers (NB), decision trees (DTs), random forests (RFs), adaptive boosting, and gradient boosting. In studies comparing multiple approaches on the same dataset, no single method consistently outperformed others. Feature extraction was predominantly automated radiomic pipelines, though a small number of studies used manual feature extraction. Only one study incorporated clinical features alongside imaging radiomic features.
Performance on GBM versus PCNSL: For the primary diagnostic application, AUC values were mainly above 0.90. Selected results illustrate the range: Alcaide-Leon et al. achieved AUC 0.877 with an SVM using 153 PET texture features, compared to radiologist AUCs of 0.845 to 0.899, demonstrating non-inferiority. Chen et al. achieved LDA-based AUCs of 0.956 to 0.978 on a validation cohort. Kim et al. reported SVM AUC 0.997 in discovery dropping to 0.947 in an independent validation cohort, with RF achieving AUC 1.00 in discovery and 0.953 in validation. Nakagawa et al. achieved AUC 0.98 using XGBoost. Zhang et al., using an end-to-end deep learning model, reported AUCs of 1.00 for GBM, 0.96 for PCNSL, and 0.954 for tumefactive demyelinating lesions in a dataset of 261 cases. Kang et al., one of only three studies using an external validation set from a different institution, reported training AUCs of 0.910 to 0.983 and external test AUCs of 0.787 to 0.946 depending on classification and feature selection method, using a heterogeneous MRI protocol that confirmed some robustness.
Performance on other diagnostic tasks: For differentiating lymphoma from other benign or malignant lesions at various sites, AUC values were generally above 0.80. Hou et al. achieved AUC 0.803 (sensitivity 71.4%, specificity 90.5%) for distinguishing idiopathic orbital inflammation from ocular adnexal lymphoma, compared with experienced radiologist sensitivity of 75.0% and specificity of 67.9%. Seidler et al. achieved AUC 0.95 with RF and 0.99 with gradient boosting for distinguishing lymphoma from normal nodes and inflammatory nodes using dual-energy CT. For disease location and detection, Sibille et al. achieved AUC 0.97, 0.84, and 0.88 for lymphoma localization to body part, organ, and subregion respectively using a CNN on FDG-PET/CT, with 0.95 classification accuracy. Li et al.'s RF model for identifying bone marrow involvement in suspected leukemia relapse achieved 87.5% sensitivity, 89.5% specificity, and 88.6% accuracy versus 62.5%, 73.7%, and 68.6% for visual inspection.
Comparisons with radiologists: In every study that compared ML performance directly with human interpretation, ML methods were reported as equivalent or superior, with one notable exception: Swinburne et al. reported maximum ML accuracy of 69.2% for the three-class problem of distinguishing GBM, PCNSL, and brain metastases, compared with 65.4% and 80.8% for two human readers. However, when the ML output was added to routine human interpretation rather than used in isolation, there was a 19% increase in diagnostic yield, illustrating the concept of human-AI complementarity rather than replacement.
Eleven of the 53 included studies applied ML exclusively to segmentation tasks, all using FDG-PET/CT images with one exception (PCNSL segmentation on MRI). The disease distributions were: three studies on DLBCL, one on both DLBCL and HL, one on DLBCL and NHL, one on NK/T-cell lymphoma, one on HL, and four on NHL or unspecified lymphoma. CNN-based methods dominated (eight studies), with random forests, adversarial networks, and conditional random fields also represented. The primary goal in most studies was computing TMTV as an automated substitute for the labor-intensive manual ROI delineation currently required in clinical practice.
Ground truth variability: A fundamental problem identified across segmentation studies is the lack of a standardized ground truth. Different studies used 41% SUVmax adaptive thresholding of lesions after manual placement (Blanc-Durand et al., Copobianco et al.), 41% SUVmax in manually placed ROIs (Grossiord et al., Weisman et al. 2020a, Yu et al.), manual segmentation by radiologists without further methodological details (Hu et al., Pennig et al., Sadik et al., Weisman et al. 2020b, Jemaa et al.), or no stated method (Yuan et al.). This variability introduces intra- and interobserver error, compromises reproducibility, and makes meaningful cross-study comparison nearly impossible.
Segmentation performance metrics: All studies reported Dice similarity coefficients (DSCs), which measure the overlap between automated and reference segmentations on a scale from 0 to 1. Values ranged from 0.71 to 0.95 across studies. Blanc-Durand et al. reported a mean DSC of 0.73 on a validation set of 94 DLBCL patients (n=639 training), with R-squared values of 0.88 and 0.82 for TMTV correlation in two cohorts. Copobianco et al. in n=280 DLBCL patients reported Dice score 0.73, significant TMTV correlation (rho=0.76, p less than 0.001), and showed that AI-derived TMTV predicted progression-free survival (PFS hazard ratio 2.4) and overall survival (OS hazard ratio 2.8), comparable to reference TMTV-based predictions (PFS HR 2.6, OS HR 3.7). Jemaa et al. achieved DSC 0.895 in training and 0.886 in testing with TMTV and SUVmax correlations of 0.97 and 0.96 with radiologist ground truth. Weisman et al. (HL, n=100 pediatric patients) reported median DSC 0.86 with Pearson's R of 0.95 and 0.88 for SUVmax and MTV versus physician contours, although MTV was slightly underestimated. Pennig et al. reported high volumetric correlation for PCNSL segmentation on MRI (TTV: r=0.88, p less than 0.0001; core: r=0.86, p less than 0.0001) with median DSC of 0.76.
Validation weaknesses: No segmentation study validated its algorithm on a fully independent cohort from a separate institution, which is the standard required for clinical translation. Studies relied on random splits, cross-validation, separate datasets from the same institution, or reported no validation at all. The QUADAS-2 assessment classified all segmentation studies as high risk of bias in the patient selection domain (due to their retrospective case-control design) and as unclear for reference standard bias and flow and timing bias, since none stated whether reference standard assessment was conducted independently of ML results.
Nine studies applied ML to prognostication or predicting responses to therapy. The disease scope was broader than the diagnostic studies: predicting overall survival or progression-free survival in extranodal NK/T-cell lymphoma nasal type (Guo et al.), multiple myeloma (Jamet et al., Morvan et al.), DLBCL (Jullien et al.), and mantle cell lymphoma (Mayerhoefer et al. 2019); predicting treatment response in DLBCL (Coskun et al., Santiago et al.) and Hodgkin lymphoma (Milgrom et al.); and identifying high-risk cytogenetic multiple myeloma patients (Liu et al.). Six studies used FDG-PET/CT, two used CT, and one used MRI.
Algorithms and features: The prognostic studies used a diverse set of ML approaches, including logistic regression classifiers, random survival forests (RSFs), weakly supervised deep learning, ANN/CNNs, SVM, RF, decision trees, K-NN, and XGBoost. A notable strength in this subgroup compared to the diagnostic studies was the integration of non-imaging variables: two studies incorporated clinicopathological variables alongside radiomic features, one added laboratory variables (LDH, WBC, Ki-67 index, ECOG performance status), and one added two additional radiological features (nodal site and subjective necrosis). This multimodal integration generally improved performance.
Performance results: Mayerhoefer et al. (n=107 mantle cell lymphoma) achieved AUC 0.83 for predicting 2-year progression-free survival, compared with AUC 0.73 for radiomic features alone, by adding LDH, WBC, Ki-67, and ECOG status to the model. Milgrom et al. (n=251 HL) achieved AUC 0.95 for predicting refractory or relapsed HL using an SVM with 5 radiomic features, compared with AUC 0.78 for MTV and TLG and 0.65 for SUVmax alone. Santiago et al. achieved AUC 0.83 and 0.79 (two readers) for predicting primary treatment failure to R-CHOP in DLBCL using RF with 1,218 handcrafted radiomic features reduced to 66, compared with AUC 0.56 and 0.52 for subjective necrosis assessment alone. Guo et al. (weakly supervised deep learning, n=64 training/n=20 test in extranodal NK/T-cell lymphoma) achieved AUC 0.99 in training and 0.88 in testing for PFS prediction. Morvan et al. (n=66 multiple myeloma from a multicenter study) reported average prediction error of 0.36 using random survival forests, compared with 0.43 to 0.61 for conventional approaches. The hazard ratios in survival-discrimination studies were HR 4.3 (Jamet et al., multiple myeloma) and approximately HR 2 (Jullien et al., DLBCL) between good and poor prognosis groups.
Quality assessment: NOS scoring classified six studies as good quality, two as fair quality, and one as poor quality, the latter because cohorts were not comparable due to inadequate control for confounders (zero stars in the comparability domain). No prognostic study validated models on external independent datasets; all used data splits or cross-validation. Critically, no study reported calibration statistics, meaning the actual probability estimates generated by the models (for example, "this patient has a 70% probability of treatment failure") were not validated against observed event rates.
The QUADAS-2 assessment yielded a sobering picture: every single diagnostic and segmentation study was classified as high risk of bias in the patient selection domain. The reason is structural and applies to the entire field: these studies are retrospective case-control designs in which patient outcomes were already known before ML algorithms were applied. This creates selection bias because cases are typically enriched with confirmed disease, the prevalence of disease in the study population does not reflect the prevalence in the target clinical population, and the heterogeneity of disease presentations seen in actual practice is underrepresented. This case-control design is essentially unavoidable in early-stage ML development where labeled datasets are needed to train models, but it means that the reported AUCs and accuracies systematically overestimate performance in prospective clinical use.
Index test and reference standard domains: Conversely, all studies were assessed as low risk of bias in the index test domain because the ground truth was not visible during computational analysis and algorithm development defined a prespecified threshold used in the test phase. The reference standard domain was more problematic: while nine studies explicitly stated that the reference was interpreted without knowledge of ML results, the remainder left this uncertain. The lack of reporting on reference standard independence also created uncertainty in the flow and timing domain, since the interval between index and reference tests was unclear in many studies.
Applicability concerns: Only five studies validated algorithms on external validation cohorts (two via temporal rather than geographic splits, which the authors note is potentially biased since the same institution's equipment and protocols are used). All other studies relied on internal random splits or cross-validation, which cannot account for training dataset biases introduced by patient selection or specific scanner characteristics. This has direct implications for applicability: a model trained on data from a single center using one MRI scanner brand or PET protocol may not perform equivalently when deployed at a different institution.
Missing calibration and performance metrics: No study in the entire scoping review, across all 53 papers, reported calibration statistics. Calibration refers to establishing the uncertainty of risk estimates or classifications, specifically whether the model's predicted probability for an event matches the actual observed frequency of that event. This is particularly important for prognostic models: a model predicting that a patient has a 70% probability of treatment failure is only clinically useful if patients predicted at 70% actually experience treatment failure approximately 70% of the time. A highly discriminatory but poorly calibrated model (which can distinguish patients who will fail from those who will not, but assigns wildly inaccurate probability values) has limited clinical utility for guiding treatment decisions. The authors note this is a field-wide problem: one reported systematic review found that 79% of 71 ML clinical prediction studies failed to address calibration.
Category 1: Model application, validation, and performance evaluation. No single ML approach was identified as consistently superior, and without head-to-head comparisons on the same datasets using standardized protocols, no meaningful comparison between methods is possible. The dominant use of cross-validation and internal data splits instead of independent external datasets is analogous to the failure of molecular biomarker development, where hundreds of candidate biomarkers identified in discovery cohorts never reached clinical practice because they were never validated in independent datasets. Cross-validation does not account for training dataset bias introduced by patient selection or specific scanner characteristics. While increasing sample size helps, the authors cite Riley et al.'s sample size framework for predictive models as a more principled approach: tailoring sample sizes to the specific clinical context and statistical requirements of the model, rather than simply collecting more data.
Category 2: Methodology and reporting standards. The authors evaluated adherence to CLAIM (Checklist for Artificial Intelligence in Medical Imaging), a reporting guideline published in 2020 specifically for ML in medical imaging. None of the included studies used or fully adhered to CLAIM criteria, even those published within the 12 months preceding the review's search cutoff. In a systematic review of CLAIM compliance in 186 ML radiology studies, median CLAIM compliance was only 0.40 (IQR 0.33 to 0.49), with only 27% documenting eligibility criteria and 49% assessing model performance on test data partitions. The authors recommend that all future studies use CLAIM from the outset, which would ensure standardized reporting of reference standards, validation approaches, performance metrics, and feature reliability assessments (i.e., testing whether the same features would be extracted with a different scanner or scanning protocol).
Category 3: Study populations and generalizability. Most study populations were small (fewer than 100 patients in the majority of studies) and heterogeneous (often mixing all grades or subtypes of lymphoma). Only two studies were multinational (Germany/USA and Denmark/USA), and only two were conducted in low- or middle-income countries (LMICs; Turkey and Tunisia). The geographic concentration of studies in China, the UK/Europe, North America, South Korea, and Japan means that models may not generalize to populations in settings with different disease biology, treatment protocols, equipment, or imaging practice patterns. Including LMICs in multicentric studies would not only improve generalizability but could also reveal geographic variation in hematological malignancy biology, generating scientifically valuable insights beyond mere validation.
Segmentation-specific gaps: For segmentation tasks, the lack of standardized TMTV definition is a separate and equally critical problem. It is currently unclear which SUV threshold is optimal (41%, 2.5, or mean liver uptake), and none of the existing segmentation studies have used a standardized ground truth. Establishing a consensus definition, standardizing volume segmentation methods, and determining guidelines for which tumor-bearing anatomical regions to include are prerequisites for an automated method that could be deployed clinically and replicated across institutions.
The authors conclude with a six-point research agenda derived directly from the identified gaps. First, adherence to CLAIM and other standardized reporting guidelines to reduce bias and improve comparability across studies. Second, validation of models in independent cohorts of sufficient size calculated a priori using frameworks such as the Riley et al. sample size methodology, rather than simply applying cross-validation or random splits within a single institution's dataset. Third, establishing a strict definition of TMTV and standardizing volume segmentation methods, including consensus on which SUV threshold to use and which anatomical regions to include, to enable both clinical adoption and valid cross-study comparisons.
Prospective study designs: Fourth, establishing comprehensive prospective studies that include different tumor grades, direct comparisons with radiologist performance, and coverage of multiple imaging modalities, sequences, and planes. Prospective designs reduce the inherent case-control bias of retrospective studies and allow the algorithm to be tested in a realistic clinical workflow where disease prevalence and presentation heterogeneity match actual practice. Fifth, comparing different ML methods head-to-head on the same cohort to definitively explore which methods perform best for specific applications and contexts, rather than each study developing a new method on its own dataset without reference to previous work. Sixth, including low- and middle-income countries in multicentric study designs to improve generalizability and reduce inequity in access to AI-enhanced diagnostics.
The role of AI in clinical decision support: The authors draw an important distinction between using AI to replace radiological assessment and using AI to facilitate clinical decision-making. For high-stakes binary decisions such as PCNSL versus GBM, an imperfect model is not necessarily useless; the 19% increase in diagnostic yield when ML output was added to human reading (Swinburne et al.) illustrates how AI can augment rather than displace expertise. For treatment decision-making, the clinical value will likely come from AI-driven radiological biomarkers that quantify features such as tumor heterogeneity or total metabolic burden in ways that are too complex or time-consuming for human assessment, providing probabilistic input to treatment algorithms rather than binary yes/no outputs.
Regulatory pathway: The authors note that ML-based software intended to diagnose, treat, or prevent disease is classified as a medical device under the Food, Drug, and Cosmetic Act (software as a medical device, SaMD). The FDA and comparable regulators have proposed frameworks for ensuring the safety and efficacy of ML-based SaMDs that require demonstrating meaningful clinical impact. Addressing the research gaps identified in this review would streamline the regulatory pathway by generating the prospective validation evidence that regulators require, while simultaneously improving real-world algorithm quality and applicability.