Combining Gene Expression Profiling and Machine Learning to Diagnose B-Cell Non-Hodgkin Lymphoma

Blood Cancer Journal 2020 AI 8 Explanations View Original
Original Paper (PDF)

Unable to display PDF. Download it here or view on PMC.

Plain-English Explanations
Pages 1-2
Why Diagnosing B-Cell Lymphomas Remains a Hard Problem

B-cell non-Hodgkin lymphomas (B-NHLs) are a heterogeneous group of mature B-cell malignancies that collectively represent some of the most diagnostically challenging tumors in all of pathology. The challenge stems from the fact that these cancers closely mimic the normal stages of B-cell differentiation, meaning malignant cells often look and behave much like the healthy cells they arose from. The WHO classification recognizes many distinct B-NHL subtypes, and accurate subtyping is clinically critical because treatment approaches differ substantially across entities, from watchful waiting for indolent low-grade lymphomas to aggressive chemoimmunotherapy for high-grade diffuse large B-cell lymphoma (DLBCL).

Current diagnostic complexity: Standard-of-care diagnosis requires integration of multiple complementary methods: morphological examination of stained tissue sections, immunohistochemical (IHC) staining panels, immunoglobulin clonality assessment, flow cytometry, conventional cytogenetics, fluorescence in situ hybridization (FISH), and increasingly next-generation DNA sequencing. Despite this multi-modal approach, the error rate in B-NHL diagnosis remains higher than in most other areas of pathology. Published studies document that expert secondary review at specialized centers changes the diagnosis in a meaningful proportion of cases, with discordance rates that underscore the genuine difficulty of the classification task.

What this paper proposes: Bobee et al. (2020), working at Centre Henri Becquerel and collaborating institutions in France, developed a middle-throughput gene expression assay that combines ligation-dependent RT-PCR with next-generation sequencing (the RT-MLPSeq platform) and applies a random forest machine learning classifier to discriminate the seven most clinically important B-NHL subtypes in a single experiment. The seven subtypes are: ABC DLBCL, GCB DLBCL, primary mediastinal B-cell lymphoma (PMBL), follicular lymphoma (FL), mantle cell lymphoma (MCL), small lymphocytic lymphoma (SLL), and marginal zone lymphoma (MZL, a category that also groups MALT lymphomas and lymphoplasmacytic lymphomas).

The assay was designed to be implementable in any molecular diagnostic laboratory already running next-generation sequencing workflows, without requiring any specialized proprietary platform. The authors explicitly positioned it as a complement to conventional histology rather than a replacement, offering a systematic readout of over 130 diagnostic and prognostic markers in parallel.

TL;DR: B-NHL diagnosis requires integrating IHC, FISH, flow cytometry, cytogenetics, and NGS, yet error rates remain high relative to other cancers. This paper presents a combined RT-PCR/NGS gene expression assay plus a random forest classifier trained on 400+ expert-annotated cases to classify seven B-NHL subtypes simultaneously. The platform is designed for routine diagnostic labs without requiring proprietary instrumentation.
Pages 2-4
The RT-MLPSeq Assay: 137 Markers, One Experiment

The RT-MLPSeq assay is built on RT-MLPA (reverse transcription multiplex ligation-dependent probe amplification) combined with next-generation sequencing. The procedure involves four core steps: reverse transcription of total RNA into cDNA, hybridization of a panel of ligation-dependent PCR probes, ligation using a thermostable DNA ligase, and PCR amplification with barcoded primers. Amplification products are purified with AMPure XP beads and sequenced on an Illumina MiSeq instrument. All reads are normalized using unique molecular identifiers (UMI), 7-nucleotide random sequences embedded in each probe, to eliminate PCR amplification bias. Samples are considered interpretable when at least 5,000 distinct UMI sequences are detected across the full marker panel.

Gene panel design: A curated panel of 137 gene expression markers was designed to cover the breadth of B-NHL biology, organized into functional groups: B-cell differentiation markers, T-cell markers (TCRalpha, TCRbeta, TCRgamma, TCRdelta, CD3, CD4, CD5, CD8, and T-cell transcription factors), macrophage and immune response markers (CD68, CD163), therapeutic targets, prognostic biomarkers, and immunoglobulin gene transcripts including sterile transcripts required for class switch recombination. The panel also detects five recurrent somatic point mutations (XPO1 E571K, MYD88 L265P, BRAF V600E, IDH2 R172K, RHOA G17V) and two viral infection markers (EBER1, HTLV1). Two probe pairs were designed for key genes such as AICDA, BCL6, MYC, and BCL2 to improve measurement accuracy.

Patient cohort: 510 B-NHL biopsies were analyzed in total: 325 DLBCL, 43 PMBL, 55 FL, 31 MCL, 17 SLL, 20 nodal or splenic MZL, 11 extranodal MALT lymphomas, and 8 lymphoplasmacytic lymphomas. The cohort combined cases from a single academic institution (Centre Henri Becquerel, Rouen) with samples from two prospective clinical trials, SENIOR (NCT02128061, n=96) and RT3 (NCT03104478, n=48). All diagnoses were established according to 2016 WHO criteria by a panel of expert pathologists from the LYSA (Lymphoma Study Association). RNA was extracted from both FFPE and fresh-frozen tissue, with extraction performed using three validated commercial systems depending on sample source.

Random forest training: The random forest classifier was trained using the scikit-learn Python library with standard hyperparameters: Gini index attribute selection, max_depth of 20, and min_samples_split of 4. The final model consists of 5,000 decision trees. The 429 cases eligible for model training (after excluding EBV-positive DLBCL, grade 3B FL, and ambiguously classified DLBCLs) were randomly split two-thirds/one-third into training (n=283) and validation (n=146) cohorts. The training cohort contained 190 DLBCLs (76 ABC, 86 GCB, 28 PMBL), 35 FLs, 21 MCLs, 12 SLLs, and 25 MZL-group cases.

TL;DR: The RT-MLPSeq assay measures 137 markers covering B-cell, T-cell, macrophage, immunoglobulin, and mutation-detection targets in a four-step PCR/NGS workflow. 510 expert-annotated B-NHL biopsies were used; 429 eligible cases split 283/146 for training/validation of a random forest with 5,000 decision trees. Technical validation used comparison against the Nanostring Lymph2Cx assay for 96 FFPE samples.
Pages 4-7
Cell-of-Origin Signatures and Separating DLBCL from PMBL

The most clinically impactful classification challenge within DLBCL is the cell-of-origin (COO) distinction between the activated B-cell (ABC) and germinal center B-cell (GCB) subtypes. This distinction matters because ABC DLBCL has a consistently worse prognosis under standard R-CHOP chemoimmunotherapy compared to GCB DLBCL, and ongoing trials are testing whether subtype-specific therapeutic modifications can close this outcome gap. The gold-standard COO assay, the Nanostring Lymph2Cx, measures 15 genes, whereas the RT-MLPSeq panel captures a far broader set of COO-related markers embedded within its 137-gene panel.

PCA and volcano plot findings: Unsupervised principal component analysis (PCA) of 125 ABC and 127 GCB DLBCL cases cleanly separated the two subtypes along the first principal components. Differential gene expression analysis confirmed the expected signatures: ABC DLBCL was characterized by high expression of TACI, FOXP1, LIMD1, IRF4, PIM2, CCDC50, CREB3L2, CYB5R2, SH3BP5, and RAB7L1, while GCB DLBCL was marked by CD10, LMO2, ASB13, NEK6, MYBL1, MAML3, ITPKB, SERPINA9, S1PR2, and BCL6. Interestingly, PCA also identified a COO-independent T-cell component (CD28, BAFF, CD3, GATA3, CD8, PRF) that likely reflects variable T-cell infiltration in the tumor microenvironment, an important confound in single-marker IHC-based classification.

Separating PMBL from other DLBCLs: Primary mediastinal B-cell lymphoma is morphologically similar to GCB DLBCL but has a distinct biology and clinical behavior, making its accurate identification important. The assay successfully retrieved the three expected signatures in PMBL versus ABC and PMBL versus GCB PCA comparisons. PMBL-specific overexpression of CD30 and CD23 was confirmed at the RNA level, consistent with routine diagnostic IHC. High expression of PDL1, PDL2, and JAK2, and the expected downregulation of BANK, CARD11, and TCL1A in PMBL were also confirmed, aligning with prior transcriptomic characterizations of this entity by Rosenwald et al. These JAK/STAT pathway markers and immune checkpoint genes are particularly important because PDL1/PDL2 overexpression in PMBL is the biological basis for checkpoint inhibitor trials now entering clinical practice.

The PCA of GCB DLBCL versus FL identified three major components. The GCB DLBCL component was dominated by canonical GCB markers (CD10, MYBL1, NEK6, BCL6) plus the KI67 proliferation marker and tumor-associated macrophage (TAM) marker CD68, along with cytotoxic (GRZB, PRF) and immune escape markers (PDL1, PDL2). The FL component was characterized by abundant T-cell markers (CD3, CD5, CD28, CTLA4, GATA3, CCR4) and follicular helper T-cell (Tfh) markers (ICOS, CD40L, CXCL13), reflecting the dense follicular microenvironment that characterizes FL tissue architecture.

TL;DR: PCA cleanly separates ABC from GCB DLBCL using expected signatures (FOXP1/IRF4/TACI for ABC; CD10/LMO2/BCL6 for GCB) plus an independent T-cell infiltration component. PMBL is distinguished by JAK/STAT activation (JAK2, PDL1, PDL2, CD30, CD23) and downregulation of BANK/CARD11/TCL1A. GCB DLBCL versus FL separation leverages Tfh markers (ICOS, CD40L, CXCL13) and proliferation/immune escape differences.
Pages 7-9
Distinguishing Indolent B-NHLs: MCL, SLL, FL, and MZL

Classifying small B-cell lymphomas, the indolent entities, is often the most challenging diagnostic task because these tumors are morphologically subtle, share many surface markers, and can be confused with one another or with reactive processes. The panel addresses this by capturing the combined expression patterns of neoplastic cells and their surrounding microenvironment, each of which contributes unique diagnostic information.

PCA of low-grade B-NHLs: Restricting the PCA to low-grade cases revealed two major separating components. The first, associated with FL, grouped GCB markers (BCL6, MYBL1, CD10, LMO2) and T-cell markers (CD28, ICOS), reflecting both the germinal center origin of the tumor cells and the characteristic Tfh-rich microenvironment of FL. The second component, grouping the other small B-cell lymphomas, was dominated by activated B-cell markers (LIMD1, TACI, SH3BP5, CCDC50, IRF4, FOXP1), consistent with the late germinal center or memory B-cell origin of MCL, SLL, and MZL.

Subtype-specific marker profiles: The assay successfully retrieved the canonical SLL phenotype: CD5-positive, CD23-positive, CD10-negative, with additional CD27 expression consistent with mature B-cell origin and elevated JAK2 expression suggesting JAK/STAT pathway activation. Importantly, SLL tumors also showed downregulated SH3BP5, a negative regulator of Bruton's tyrosine kinase (BTK), pointing to a potential mechanism for constitutive BTK signaling in this disease. For MCL, the assay correctly identified the CCND1-high, CD5-high, BCL2-high phenotype along with expected downregulation of CD10 and CD23. TCL1A and CCDC50, both prognostic in MCL, were overexpressed, as was the B-cell chemokine receptor CXCR5, which is involved in tumor dissemination. The MZL group showed the expected triple-negative phenotype (CD5-negative, CD10-negative, CD23-negative) with high CD138 expression and low KI67, consistent with indolent biology.

The PCA and differential expression analysis of DLBCL versus the pooled small cell lymphomas confirmed that high-grade aggressive lymphomas across all COO subtypes share a set of common features: elevated KI67, high germinal center-associated gene expression, TAM markers (CD68, CD163), cytotoxic markers (GRZB, PRF), and immune checkpoint upregulation (PDL1, PDL2). Low-grade lymphomas were instead characterized by T-cell markers and the follicular dendritic cell marker CD23, reflecting their dependence on microenvironmental survival signals.

TL;DR: SLL is identified by CD5+/CD23+/CD10-/JAK2-high/SH3BP5-low. MCL by CCND1-high/CD5-high/BCL2-high/TCL1A-high/CXCR5-high. MZL by triple-negative (CD5-/CD10-/CD23-)/CD138-high/KI67-low. High-grade versus low-grade lymphoma segregation is driven by KI67, TAM markers, cytotoxic markers, and immune checkpoints versus T-cell and follicular dendritic cell markers.
Pages 8-10
IGH Transcript Patterns as Diagnostic Discriminators

Beyond cellular gene expression, B-cell NHLs differ in their immunoglobulin (Ig) gene configurations and transcriptional states, reflecting their B-cell maturation stage and isotype status. The RT-MLPSeq panel includes a comprehensive set of immunoglobulin gene probes covering IGHM, IGHD, and all sterile transcripts required for class switch recombination (CSR), making it possible to assess Ig gene expression patterns simultaneously with cell lineage and functional markers.

IGHM and IGHD expression patterns: MCL and SLL can be distinguished from other B-NHLs by their high expression of the IGHD gene, consistent with their derivation from naive or minimally activated B cells that have not yet undergone CSR. Two broad groups of tumors were defined based on IGHM expression. The first, IGHM-positive, encompasses tumors with activated or memory B-cell origin: most ABC DLBCLs, MCL, MZL, and SLL. The second, IGHM-negative, includes GCB-origin tumors (GCB DLBCL and FL), which have typically completed isotype switching, and PMBLs, which generally lack Ig expression altogether.

Class switch recombination defect in ABC DLBCL: A particularly informative finding involves the Imu-Cmu sterile transcript, which controls the accessibility of the switch-mu region to the CSR machinery. In normal IgM-positive B cells that are undergoing CSR, AICDA (activation-induced cytidine deaminase) expression is accompanied by activation of the appropriate sterile transcript (Imu-Cmu for switching away from IgM). The data confirm a CSR defect in ABC DLBCL: these tumors paradoxically express AICDA together with the IGHM gene, yet the Imu-Cmu transcript is specifically downregulated, apparently preventing isotype switching despite AICDA activity. In contrast, IgM-positive NHLs without CSR defect (SLL, MZL, MCL) express both IGHM and Imu-Cmu as expected.

An additional striking finding is the near-exclusive expression of the Iepsilon-Cepsilon sterile transcript in FL samples, making it one of the most discriminatory markers for that entity in the entire panel. The Igamma-Cgamma transcript was also found to be highly expressed in SLL and MCL, two non-germinal-center-derived lymphomas, an unexpected finding that warrants further biological investigation. These Ig transcript patterns add a layer of discriminatory information that is not captured by conventional surface marker IHC panels.

TL;DR: IGHD high expression marks MCL and SLL. ABC DLBCL has a CSR defect: AICDA is expressed but Imu-Cmu is downregulated, blocking isotype switching despite IGHM expression. The Iepsilon-Cepsilon sterile transcript is near-exclusively expressed in FL, making it a powerful discriminatory marker. Ig transcript patterns captured by the panel provide diagnostic information not available from standard IHC.
Pages 9-11
Validation Results: 94.5% Accuracy Across Seven Subtypes

The random forest (RF) classifier was trained on the 283-case training cohort and independently validated on the 146-case validation cohort. In the training cohort, the RF algorithm classified all 283 cases into the expected subtype with high-confidence probability distributions showing clear separation between the predicted class and all other classes. These training results, while expected given model fitting, established the discriminability of the feature space.

Validation cohort performance: In the independent validation cohort, the RF predictor correctly classified 138 of 146 cases (94.5%), demonstrating strong generalization. For ABC and GCB DLBCL, the concordance with the Lymph2Cx NanoString assay was 94.3%. Importantly, all 49 ABC DLBCLs in the validation cohort were correctly classified (100% concordance), while 36 of 41 GCB DLBCLs were correctly classified (87.8%). The five GCB discordances were informative: two cases classified as PMBL by the RF predictor were subsequently found to carry mutations compatible with PMBL diagnosis (B2M, TNFRSF14, SOX11, and CIITA mutations in one case; STAT6, B2M, CD58, CIITA, and CARD11 mutations in the other), suggesting the RF predictor may have captured true biological PMBL features that the Lymph2Cx assay did not detect.

Small cell lymphoma accuracy: 14 of 15 PMBLs (93.3%) and 39 of 41 small cell lymphomas (95.1%) were accurately classified in the validation cohort, including all MCLs and all SLLs at 100% accuracy. One MZL was misclassified as FL, attributable to its preeminent GCB gene expression signature, and one FL was misclassified as GCB DLBCL in a patient presenting with stage IV disease and a leukemic presentation, where the high tumor-to-microenvironment ratio likely compressed the FL-specific microenvironmental signals the classifier depends on.

Additional analyses of cases excluded from model building provided useful insights. Five of eight FL grade 3B tumors, excluded from training due to their biologically ambiguous nature, were classified as DLBCL by the RF predictor (3 GCB, 2 ABC), while three were classified as FL, consistent with the known biological overlap of these entities. Of six DLBCLs that the Lymph2Cx assay returned as unclassified, five were classified as ABC DLBCLs by the RF predictor, including two with CD79B mutations (typically an ABC signature), and one was classified as GCB.

TL;DR: The random forest classifier achieved 94.5% overall accuracy (138/146) on independent validation. ABC DLBCL concordance with Lymph2Cx was 100% (49/49); GCB DLBCL 87.8% (36/41). All MCLs and SLLs were correctly classified. Two apparent misclassifications vs. Lymph2Cx were later found to carry genomic mutations supporting the RF predictor's PMBL call.
Pages 11-13
MYC/BCL2 Double Expression and Survival in DLBCL

Beyond classification, the assay was evaluated for its ability to capture clinically important prognostic markers in DLBCL. The survival analysis focused on 104 DLBCL patients treated with rituximab plus chemotherapy (R-CHOP or equivalent) at Centre Henri Becquerel between 2000 and 2017, with overall survival (OS) and progression-free survival (PFS) computed from treatment start and right-censored at 5 years or last follow-up. Kaplan-Meier curves were estimated for key subgroups, and significance tested with the log-rank test. Multivariable Cox proportional hazards models adjusted for IPI score and cell-of-origin classification were used to confirm independent prognostic value.

COO and IPI status: ABC cell-of-origin was significantly associated with inferior OS (p=0.0236) but showed only a trend for PFS (p=0.0699). IPI score 4-5 versus 0-3 was strongly associated with both worse PFS (p=1x10-5) and OS (p=1x10-4). MYC expression alone was associated with inferior PFS (p=5x10-5) and OS (p=4x10-4). BCL2 expression alone was associated with inferior PFS (p=1x10-3) and OS (p=5x10-3).

MYC/BCL2 double expressors: The combination of high MYC and high BCL2 expression identified a group of 25 double-positive cases (24% of the cohort) with a particularly poor prognosis. In univariate analysis, double-expressor status was significantly associated with inferior PFS (p=1x10-5) and OS (p=1x10-5). In multivariable analysis adjusted for both IPI and COO, MYC/BCL2 double expression remained independently prognostic for both OS (HR 2.08, 95% CI 1.34-3.25, p less than 5x10-3) and PFS (HR 2.04, 95% CI 1.35-3.12, p less than 5x10-3). The prognostic impact of double expression was also independent when adjusting for IPI score alone (OS HR 2.20, 95% CI 1.41-3.41; PFS HR 1.92, 95% CI 1.27-2.89). Double-expressor patients showed significant clinical correlations with higher age (p=5x10-3), elevated LDH (p=0.04), and ABC subtype (p less than 10-4).

Additional prognostic genes identified in supplemental analyses included CARD11 (PFS p less than 10-3, OS p less than 10-4), CREB3L2 (PFS p less than 10-4, OS p less than 10-4), STAT6 (PFS p less than 10-3, OS p less than 10-2), and CD30 (PFS p less than 10-2, OS p less than 10-3). These findings suggest that the broader panel captures prognostic information extending well beyond MYC and BCL2 alone, potentially identifying additional targetable pathways in high-risk DLBCL.

TL;DR: MYC/BCL2 double expressors (24% of cohort) had dramatically inferior survival (HR 2.08 OS, 2.04 PFS vs. others), independent of IPI and COO in multivariate analysis. IPI and ABC subtype also significantly predicted outcomes. Additional prognostic genes included CARD11, CREB3L2, STAT6, and CD30, all highly significant for both PFS and OS.
Pages 12-13
Clinical Potential, Limitations, and Path Forward

The authors discuss both the strengths and the constraints of the RT-MLPSeq approach in the context of routine clinical implementation. The assay's principal strengths are its wide subtype coverage (seven entities in a single experiment), its compatibility with FFPE biopsy material, its use of standard NGS equipment already deployed in many diagnostic molecular labs, and its capacity to simultaneously capture diagnostic classification and prognostic biomarker readouts from the same experiment. The random forest output also returns calibrated probability scores for each of the seven classes rather than a hard label, which provides pathologists with a confidence measure that can guide interpretation, particularly for borderline cases.

Key limitations: The authors are explicit that a rigorous initial histological evaluation of the biopsy remains mandatory before deploying the gene expression assay. The classifier was trained and validated on samples already presumed to represent B-NHL, and the assay cannot distinguish a reactive lymph node or a T-cell lymphoma from a B-NHL without this prerequisite morphological assessment. One clinically important limitation is the need for a minimum RNA concentration of 20 ng/ul and at least 5,000 UMI sequences for interpretable results, which may exclude samples from very small or poorly preserved biopsies that are common in clinical practice.

MYC/BCL2 and IHC standardization: The authors highlight a specific area where gene expression may offer advantages over standard IHC for prognostic assessment. The optimal cut-off values for defining MYC-positive and BCL2-positive cases by immunohistochemistry remain debated, with published prevalence of double-expressor DLBCL ranging from 21% to 31% across studies. The gene expression readout provides a continuous quantitative measurement that is not subject to staining variability, antibody clone differences, or pathologist scoring variability. At 24% double-expressors in their cohort, the authors' results fall within the published range, with an independent prognostic impact validated in multivariable analysis.

Looking forward, the authors propose that the assay could be deployed alongside routine NGS DNA sequencing in molecular diagnostic labs, since both workflows use the same sequencer and can be run in parallel. They also note that as new molecular subgroups of DLBCL are defined by comprehensive genomic analyses (such as those identifying C1-C5 or cluster-based subtypes), the gene expression panel could be extended to incorporate markers relevant to these emerging classifications, facilitating patient stratification into targeted therapeutic trials. Additional FISH testing remains mandatory for identifying chromosomal rearrangements of MYC and BCL2 that define true double-hit lymphomas in the high-grade B-cell lymphoma category.

TL;DR: Key limitations include dependence on prior morphological diagnosis to exclude non-B-NHL entities, minimum RNA quality requirements, and the need for complementary FISH for chromosomal rearrangements. Gene expression-based MYC/BCL2 scoring avoids IHC variability and standardization issues. The assay is designed to run alongside existing NGS DNA sequencing workflows without additional instrument investment.