High-grade B-cell lymphoma (HGBL) is a relatively new disease category introduced by the 2016 WHO classification of lymphoid neoplasms. It sits within the broader family of aggressive mature B-cell lymphomas but is distinguished from standard diffuse large B-cell lymphoma (DLBCL) by the presence of chromosomal rearrangements involving the MYC, Bcl-2, and/or Bcl-6 oncogenes. Tumors with MYC plus Bcl-2 or Bcl-6 rearrangements are called double-hit lymphoma (HGBL-DH), while those with all three rearrangements are triple-hit (HGBL-TH). A third subgroup, HGBL-NOS (not otherwise specified), shows blastoid or intermediate morphology between DLBCL and Burkitt lymphoma without the defining gene rearrangements.
The diagnostic bottleneck: Definitive HGBL diagnosis currently requires fluorescence in situ hybridization (FISH) analysis on excised biopsy tissue, which detects the specific chromosomal translocations at the MYC (8q24), Bcl-2 (18q21), and Bcl-6 (3q27) loci. FISH is expensive, labor-intensive, requires specialized equipment and expertise, and is not universally available. Importantly, no international consensus exists on which DLBCL patients should be automatically referred for FISH testing, creating clinical uncertainty about how broadly to apply the test. This makes HGBL a clear candidate for clinical decision-support tools that can flag which patients need FISH workup.
Complicating the picture: A related entity called "double-expressor lymphoma" (DEL) involves overexpression of MYC and Bcl-2 proteins by immunohistochemistry (IHC) without the underlying gene rearrangements that define HGBL-DH. DEL is more common than HGBL-DH and carries a worse prognosis than standard DLBCL, but is generally less aggressive than HGBL-DH. The overlap between DEL and HGBL creates diagnostic ambiguity because IHC alone cannot substitute for FISH in identifying true gene-rearranged HGBL. Neither Ki-67 proliferation index nor double-expressor status is independently sensitive enough to distinguish HGBL-DH from non-HGBL.
This 2022 study from Fujian Medical University Union Hospital enrolled 187 newly diagnosed B-cell lymphoma patients to systematically compare HGBL and non-HGBL patients, identify independent clinical predictors of HGBL, and then build machine learning classification models capable of guiding which patients should undergo FISH testing - all using routinely available clinical, laboratory, imaging, and pathological variables rather than specialized molecular assays.
The study retrospectively enrolled 187 patients with aggressive mature B-cell lymphomas diagnosed and followed up at Fujian Medical University Union Hospital between April 1, 2018, and April 1, 2022. The cohort included 152 cases of DLBCL and 35 cases of HGBL (HGBL-DH/TH/NOS). The overall sample was 56.1% male, 43.9% female, with an average age of 55.50 (+/- 15.02) years. All patients had at least one high-risk factor at presentation, including but not limited to advanced Ann Arbor stage, extranodal involvement, double-expressor status, or high International Prognostic Index (IPI) risk group.
Laboratory and pathological workup: FISH analysis for MYC, Bcl-2, and Bcl-6 rearrangements was performed using dual-color Break Apart rearrangement probes (Vysis LSI probes, Abbott Laboratories) on formalin-fixed paraffin-embedded (FFPE) or fresh-frozen tissue with the ThermoBrite FISH slide processing system. IHC staining for CD10, MUM-1, Bcl-2, c-MYC, Bcl-6, and Ki-67 was performed using the Ventana Benchmark ULTRA module with validated Roche monoclonal antibodies. Cell of origin (COO) was determined using the Hans algorithm, classifying cases as germinal center B-cell (GCB) or non-GCB phenotype based on CD10, Bcl-6, and MUM-1 expression by IHC. DEL was defined by MYC positivity of at least 40% and Bcl-2 positivity of at least 50%. PET/CT-derived whole-body maximum standardized uptake (SUVmax) was recorded at baseline before the first induction chemotherapy.
Feature collection and LASSO reduction: A total of 37 clinical features were collected per patient, spanning demographics, laboratory values (WBC, LDH, beta-2 microglobulin), staging (Ann Arbor), B symptoms, IPI score, extranodal involvement sites (including bone marrow, gastrointestinal tract, and CNS), histomorphology, cytogenetic complexity, IHC phenotype, DEL status, Ki-67 proliferation index, SUVmax, and EBER (Epstein-Barr virus-encoded RNA) positivity. To reduce dimensionality, LASSO (least absolute shrinkage and selection operator) binary logistic regression was applied, selecting 8 potential predictors with non-zero coefficients from the original 37 features.
Machine learning algorithms tested: Thirteen different classification algorithms were evaluated, including Gradient Boosting Classifier, CatBoost, Random Forest, Extra Trees, Extreme Gradient Boosting (XGBoost), Logistic Regression, Decision Tree, Ridge Classifier, AdaBoost, K-Nearest Neighbors, SVM-Linear Kernel, Naive Bayes, and Quadratic Discriminant Analysis. The 187 cases were split 70/30 into training and test sets. Model performance was evaluated using area under the ROC curve (AUC), confusion matrix, precision, recall, and F1 score. Statistical analysis used SPSS v26.0, Python v3.9.0, and R 4.1.1.
Among the 35 HGBL cases identified by FISH, the most common subtype was MYC/Bcl-6 HGBL-DH with 21 cases (60.0%), followed by HGBL-NOS with 7 cases (20%), HGBL-TH with 4 cases (11.4%), and MYC/Bcl-2 HGBL-DH with only 3 cases (8.6%). This distribution is notably different from European and North American series, where MYC/Bcl-2 HGBL-DH typically predominates. The authors attribute this to possible geographic differences in HGBL subtype frequency, consistent with prior observations from southern China and Taiwan where MYC/Bcl-6 HGBL-DH predominates.
Statistically significant distinguishing features: Comparing HGBL against non-HGBL, several clinical and pathological variables showed significant differences. High-grade histomorphology (necrosis, massive mitoses, or "starry sky" pattern) was more common in HGBL (p = 0.009). Bone marrow involvement was seen in 28.6% of HGBL vs. 11.8% of non-HGBL patients (p = 0.012). More than one extranodal site of involvement (E greater than 1) was present in 42.9% of HGBL vs. 23.0% of non-HGBL cases (p = 0.017). Baseline SUVmax was significantly higher in HGBL patients (p = 0.045), consistent with greater metabolic activity in these more aggressive tumors.
IHC findings: MUM-1 protein expression was negatively correlated with HGBL (p = 0.019), with significantly lower expression in the HGBL group than the non-HGBL group. This is relevant because MUM-1 is a marker of the non-GCB (activated B-cell) phenotype, and its lower expression in HGBL is consistent with the finding that many HGBL-DH cases in this cohort had the GCB phenotype, particularly the MYC/Bcl-2 subtype. By contrast, WBC elevation, b2-microglobulin elevation, CD10 positivity, Ki-67 index, EBER status, and serum LDH level did not reach statistical significance between HGBL and non-HGBL groups in univariate comparisons, though LDH trended higher in HGBL (650 vs. 502 U/L, p = 0.587).
Among HGBL subtypes, bone marrow involvement was significantly more frequent in HGBL-DH/TH than in non-HGBL-DH/TH, and c-MYC protein expression by IHC was significantly higher in HGBL-DH/TH. Chromosome karyotypes were available for 18 of 35 HGBL patients and 69 of 152 non-HGBL patients; complex karyotypes (cytogenetic complexity score greater than 2) were found in 2 HGBL and 7 non-HGBL cases with available data, limiting conclusions about karyotypic complexity.
Kaplan-Meier analysis of overall survival (OS) and progression-free survival (PFS) demonstrated significantly worse outcomes in the HGBL group compared to non-HGBL (OS difference p = 0.015). The median PFS was 280 days for HGBL patients versus 567 days for non-HGBL patients and 490 days across all patients combined. The median OS was not reached in any group during the follow-up period, but the survival curves diverged clearly and early, reflecting the rapidly aggressive course of HGBL. This substantial gap in PFS - roughly half the progression-free time compared to non-HGBL - highlights the clinical urgency of accurate upfront identification.
First-line R-CHOP response: Of the 187 patients, 17 HGBL and 79 non-HGBL patients received R-CHOP (rituximab, cyclophosphamide, doxorubicin, vincristine, and prednisone) as first-line induction therapy. The objective response rate (ORR) was 64.7% for HGBL vs. 83.5% for non-HGBL (p = 0.077), a clinically meaningful but statistically borderline difference reflecting the small HGBL sample size. The R-CHOP regimen that achieves satisfactory outcomes in most DLBCL patients is considerably less effective in HGBL, consistent with prior reports.
Duration of remission: Among the 72 patients who achieved first complete remission (CR), 6 HGBL cases and 17 non-HGBL cases subsequently relapsed. Duration of remission was significantly shorter in HGBL, with a median of 121 days versus 258 days in non-HGBL (p = 0.007). This nearly 2-fold shorter remission duration underscores the high relapse risk even among HGBL patients who initially respond to therapy. Second-line regimens used in this cohort included R-DA-EPOCH, R-DA-EDOCH, R-HyperCVAD, and R-CHOP augmented with lenalidomide, chidamide, or zanubrutinib, though the study was not designed to compare these salvage approaches statistically.
The authors noted initial curative effects in patients receiving R-CHOP combined with agents such as lenalidomide, ibrutinib, or chidamide (R-CHOP + X regimens), but these findings require validation in larger prospective cohorts. The overall conclusion is clear: R-CHOP alone is insufficient for HGBL and a novel induction regimen remains an unmet clinical need for this patient population.
LASSO regression applied to all 37 clinical features selected 8 variables with non-zero coefficients as the most predictive for HGBL status. These 8 variables, along with other clinically meaningful factors identified in related literature, were then fed into the full suite of 13 machine learning classifiers for comparative evaluation. Among all models tested, Extreme Gradient Boosting (XGBoost) achieved the best AUC, precision, and recall, with micro-average and macro-average AUC values of 0.81 and 0.70, respectively. Random Forest had the highest overall accuracy among all tested algorithms.
Choosing logistic regression for clinical deployment: Despite XGBoost outperforming logistic regression on AUC metrics, the authors ultimately selected a logistic binary regression model for the primary HGBL prediction tool. The rationale was interpretability: logistic regression produces coefficient-based odds ratios that clinicians can directly evaluate and apply, whereas tree-based ensemble methods function as relative black boxes. The logistic model's AUC and precision were considered reliable enough for clinical use, even at a modest performance tradeoff.
Independent predictors from the logistic model: The logistic regression identified four statistically significant independent predictors of HGBL: high-grade histomorphology appearance (p = 0.012, OR = 14.77), advanced Ann Arbor stage (p = 0.007, OR = 0.25), LDH above the upper limit of normal (p = 0.040, OR = 0.998), and IPI risk group 3 or 4 (p = 0.003, OR = 28.24). The model achieved micro-average and macro-average AUC values of 0.85 and 0.53 in the test set, with high predictive efficiency for non-HGBL cases but more limited sensitivity for the HGBL class itself, partly a reflection of the class imbalance (152 non-HGBL vs. 35 HGBL).
XGBoost feature importance ranking: The XGBoost feature importance plot ranked variables in descending order of predictive contribution as follows: high-grade histomorphology, c-MYC greater than 0.575, extranodal involvement greater than 1, WBC above ULN, SUVmax greater than 25.25, IPI risk group 3 or 4, male sex, age greater than 60, LDH above ULN, and Bcl-6 overexpression positivity. This ordering provides a clinically interpretable checklist of factors that, when aggregated, should raise suspicion for HGBL and prompt FISH evaluation.
Beyond the general HGBL model, the authors constructed a logistic binary regression model specifically targeting HGBL-DH prediction. This model identified patients with high-grade histomorphology, SUVmax greater than 25.25, IPI risk group 3 or 4, c-MYC overexpression greater than 0.575, extranodal involvement greater than 1 site, WBC above ULN, and LDH above ULN as more likely to carry the double-hit genotype. In the test set evaluation, the micro-average and macro-average AUC values for this HGBL-DH model were 0.61 and 0.86, respectively - notably the macro-average AUC outperformed the micro-average, reflecting better class-specific discrimination when considered individually.
MYC rearrangement prediction model: Given the particular prognostic weight of MYC rearrangement (present in all HGBL-DH/TH by definition), a separate model was built to predict MYC rearrangement status specifically. Comparing results across the multiple classifier algorithms, the Extreme Gradient Boosting approach again yielded the highest AUC for this task. The feature importance plot for MYC rearrangement prediction followed a similar ranking to the HGBL-DH model, with c-MYC IHC expression, SUVmax, and IPI risk group among the top contributors.
Clinical implication: The practical value of these subtype-specific models is to help clinicians prioritize FISH testing. Not every patient suspected of having HGBL will turn out to have HGBL-DH versus HGBL-NOS, and the therapeutic implications differ. A model that can further stratify the probability of the DH genotype from the probability of HGBL broadly could refine clinical decision-making about the urgency and scope of molecular workup. The finding that c-MYC protein overexpression by IHC (threshold greater than 0.575) features prominently in both models is consistent with prior evidence that IHC-based c-MYC overexpression, while imperfect, is a useful pre-screening marker for underlying MYC rearrangement.
The authors acknowledge that a FISH protocol limited to cases with certain high-risk clinical features (as identified by their models) would save time and cost but would inevitably miss some HGBL cases. This tradeoff between sensitivity and resource efficiency is a central tension in HGBL workup strategies, and these models are positioned as decision aids rather than replacements for clinical judgment.
The most clinically actionable model built in this study was a logistic binary regression predicting 1-year survival (alive vs. dead at 1 year from diagnosis). In the test set, the macro-average and micro-average AUC values of the ROC curve were 0.82 and 0.73, respectively. The model's validity for identifying patients at high risk of death within 1 year was notable: precision was 0.714, recall was 0.833, and F1 value was 0.769 in the test set. These metrics indicate that the model correctly identified approximately 83% of patients who would die within 1 year (high recall) while achieving reasonable precision in its positive predictions.
Predictors of 1-year mortality: The feature importance plot from the 1-year survival model identified ten variables associated with higher risk of death within 1 year: IPI risk group 3 or 4, CD10 positivity, extranodal involvement, LDH above ULN, WBC above ULN, bone marrow involvement, age greater than 60, advanced Ann Arbor stage, and SUVmax greater than 25.25. This overlapping set of predictors with the HGBL diagnostic models reinforces the concept that the same biological aggressiveness markers that flag a patient as likely HGBL also predict their near-term survival risk.
Integrated prognostic utility: The 1-year survival model can function independently of the HGBL diagnostic model, providing prognostic stratification for all BCL patients regardless of FISH result availability. This is clinically important because FISH results may take days to weeks to return, whereas survival risk estimation based on baseline clinical data can inform immediate treatment decisions such as whether to pursue clinical trial enrollment, consider stem cell transplant planning, or intensify upfront therapy.
Together, the three models (HGBL prediction, HGBL-DH prediction, and 1-year survival prediction) form an integrated decision framework: the first two guide FISH testing triage, while the third informs prognostic counseling and treatment intensity regardless of genotyping results. The authors position this framework as a practical complement to standard pathological workup in centers where access to FISH is limited or delayed.
Single-center retrospective design: All 187 patients were drawn from a single institution in Fujian Province, China, which introduces spectrum bias and limits generalizability. Patients referred to a tertiary hematology center tend to be more severely ill and more likely to have already undergone extensive workup, which may overrepresent high-risk clinical presentations in both the HGBL and non-HGBL groups. The retrospective nature means treatment allocation was not randomized and follow-up was not standardized across patients.
Class imbalance and small HGBL sample size: With only 35 HGBL cases out of 187 total (18.7%), the training data is substantially imbalanced. The logistic regression model showed high predictive efficiency for the majority non-HGBL class but more limited sensitivity specifically for HGBL, reflected in the macro-average AUC of 0.53 for the HGBL model (compared to micro-average 0.85). Class imbalance is a well-recognized challenge in rare-disease ML classification, and techniques such as SMOTE oversampling, cost-sensitive learning, or threshold adjustment were not explicitly described as having been applied.
Geographic specificity: The unusual predominance of MYC/Bcl-6 HGBL-DH (21 of 24 DH cases, 60% of all HGBL) versus the MYC/Bcl-2 predominance in Western cohorts means that the clinical features and IHC patterns driving model predictions in this dataset may not generalize to European or North American populations. Particularly, the relatively lower frequency of CD10-positive GCB phenotype cases among the DH cohort (since MYC/Bcl-6 DH is more often non-GCB) affects which IHC variables are most discriminating across different geographic settings.
Interpretability and clinical deployment considerations: While logistic regression was chosen for its interpretability, the model equation involves 15 input variables with variable-specific regression coefficients (B) and Wald statistics. Implementing this in routine clinical practice requires either a computational tool (calculator app or EHR-integrated decision support) or extensive clinician training. The authors note that multicenter studies are needed to validate and refine these models before broad deployment, and that the 1-year survival model in particular warrants prospective validation as a real-world prognostic tool. The fundamental clinical value proposition - using routine baseline data to guide selective FISH ordering - remains compelling and is directly actionable even with existing model limitations.