Deep-Learning 18F-FDG Uptake Classification Enables Total Metabolic Tumor Volume Automation in DLBCL

Journal of Nuclear Medicine 2021 AI 8 Explanations View Original
Original Paper (PDF)

Unable to display PDF. Download it here or view on PMC.

Plain-English Explanations
Pages 1-2
Why Automating TMTV Matters for DLBCL Prognosis

Diffuse large B-cell lymphoma (DLBCL) is the most common non-Hodgkin lymphoma subtype, accounting for roughly 30-40% of all NHL cases worldwide. Despite the curative potential of standard first-line chemoimmunotherapy with R-CHOP (rituximab, cyclophosphamide, doxorubicin, vincristine, and prednisone), more than 30% of patients are refractory or relapse after treatment and face poor outcomes. Early identification of high-risk patients who might benefit from treatment intensification or novel therapies is therefore a clinical priority.

TMTV as a prognostic biomarker: Total metabolic tumor volume (TMTV) derived from baseline 18F-FDG PET/CT scans has emerged as a promising alternative to conventional prognostic indices such as the International Prognostic Index (IPI), Revised IPI, and NCCN-IPI. Unlike these indices, which rely on indirect surrogates of tumor burden (age, LDH, performance status, stage, extranodal sites), TMTV directly quantifies the total volume of FDG-avid tumor across the entire body at diagnosis. Multiple studies have demonstrated that high baseline TMTV is independently associated with inferior progression-free survival (PFS) and overall survival (OS) in DLBCL.

The measurement problem: Despite its prognostic value, TMTV is not yet routinely used in clinical practice. A key barrier is the absence of a standardized, efficient measurement approach. Calculating TMTV requires segmenting all metabolically active malignant foci throughout the body on a whole-body PET/CT scan, an inherently labor-intensive task that requires expert nuclear medicine readers. Several methods have been proposed (threshold-based, region-growing, SUV-based), and they yield different absolute volume values and different optimal cutoffs for risk stratification, contributing to a lack of consensus in the literature.

The AI opportunity: This study examines whether a convolutional neural network (CNN)-based automated method called the PET Assisted Reporting System (PARS) can estimate TMTV with prognostic performance comparable to that of expert human readers, without requiring the extensive manual input that currently limits clinical adoption. The approach combines an automated whole-body high-uptake region segmentation algorithm with a deep-learning classifier that distinguishes physiologic from tumor-related FDG uptake.

TL;DR: DLBCL accounts for 30-40% of NHL cases and over 30% of patients relapse after R-CHOP. Baseline TMTV from FDG PET/CT is a validated prognostic marker but requires time-consuming expert segmentation of all tumor sites body-wide. This study tests a CNN-based system (PARS) that automates the full workflow of detecting and classifying FDG-avid regions to produce TMTV without manual input.
Pages 2-3
The REMARC Trial Cohort and Image Analysis Setup

The study retrospectively analyzed 280 DLBCL patients from an ancillary study of the REMARC trial (NCT01122472), a phase III randomized controlled trial evaluating lenalidomide versus placebo maintenance therapy in elderly DLBCL patients (aged 60-80 years) who had responded to standard first-line R-CHOP. Patients were enrolled from 124 centers across multiple countries, making this one of the larger and more geographically diverse cohorts used to evaluate automated TMTV methods. All patients underwent baseline 18F-FDG PET/CT before initiating R-CHOP, and TMTV had previously been shown to be a strong predictor of 4-year PFS and OS in this cohort.

Patient characteristics: The cohort was 57.5% male with a median age of 68 years (range 58-80). Ann Arbor staging was predominantly advanced, with 70.4% of patients presenting at stage IV and 20.4% at stage III. The International Prognostic Index score distribution reflected a high-risk population: 34.6% had IPI score 3 and 28.9% had IPI score 4. Elevated lactate dehydrogenase (LDH), an adverse prognostic marker, was present in 58.9% of patients. After a median follow-up of 5 years, 86 patients had experienced a PFS event and 51 patients had an OS event, yielding 4-year survival rates of 69% for PFS and 83% for OS.

Image acquisition diversity: A key strength of this dataset was scanner heterogeneity. PET/CT images were acquired on different scanner models from multiple vendors across 124 centers, with variable reconstruction settings. The delay between 18F-FDG injection and image acquisition averaged 71.7 minutes (SD 14.1 min). Images were collected in anonymized DICOM format, and patients with incomplete axial slices or irregular slice intervals were excluded, leaving 280 of the 301 enrolled patients eligible for analysis.

Reference TMTV measurement: Two experienced nuclear medicine physicians independently measured the reference TMTV (TMTVREF) using a semiautomatic version of the Beth Israel Fiji (ImageJ) plugin, a validated tool previously used in multiple lymphoma TMTV studies. The workflow combined automated region detection (using a component-tree and shape-prior algorithm followed by region growing and 41% SUVmax thresholding) with manual review by the reader, who removed non-lymphoma regions and added any missed tumor sites by drawing prisms around lesions and applying the same 41% SUVmax threshold.

TL;DR: 280 DLBCL patients from the international REMARC trial (124 centers) provided baseline PET/CT scans. The cohort was predominantly advanced-stage (70.4% stage IV) with median age 68. Reference TMTV was measured by 2 expert readers using the Beth Israel Fiji plugin with combined automated and manual steps. Median follow-up was 5 years with 4-year PFS of 69% and OS of 83%.
Pages 2-3
How PARS Automates PET/CT Region Detection and Classification

The PARS (PET Assisted Reporting System) prototype operates in two sequential automated stages. In the first stage, a multi-foci segmentation (MFS) algorithm scans the entire body PET/CT for regions with elevated tracer uptake. Following PERCIST (PET Response Criteria in Solid Tumors) recommendations, the algorithm first identifies a cylindric reference region at the center of the proximal descending aorta using a CT landmarking algorithm. This region provides the mean blood pool SUV and its standard deviation. Only PET regions with an SUVpeak greater than twice the mean blood pool SUV plus twice the blood pool SUV standard deviation are considered for further analysis, yielding an average SUVpeak threshold of 3.6 (SD 1.2) across the cohort. Detected regions are further refined using 42% of SUVmax thresholding, and those with volumes below 2 cm3 are discarded to reduce false-positive noise.

The CNN classification stage: The regions of interest (ROIPARS) identified by MFS are then passed to a CNN classifier. This CNN was not developed specifically for DLBCL or TMTV computation; it was originally trained on a separate cohort of over 600 patients with lymphoma (multiple subtypes) and lung cancer undergoing PET/CT for staging and response assessment. For each ROIPARS, the CNN outputs two pieces of information: the anatomic localization of the region among a predefined set of relevant staging sites, and a binary classification of the uptake as either physiologic (e.g., bowel activity, muscle activation, inflammation, infection, or degenerative changes) or suspicious (i.e., likely lymphoma).

TMTV calculation: The PARS-based TMTV (TMTVPARS) is computed simply as the sum of the volumes of all ROIPARS sites classified as suspicious by the CNN. No manual correction or reader input is required at any step. The researchers also tested two alternative settings for the initial high-uptake ROI segmentation: one using a fixed SUV threshold of 2.5 instead of the blood-pool-based PERCIST threshold, and one that included smaller ROIs (volume 0.1-2 cm3) that were excluded in the primary analysis.

Statistical framework: Uptake classification performance was evaluated using standard sensitivity, specificity, and accuracy metrics, comparing CNN labels to the expert reference at both region and patient levels. TMTV agreement was assessed with Bland-Altman analysis and Spearman rank correlation. Survival analyses used receiver-operating-characteristic (ROC) curves to determine optimal TMTV cutoffs (maximizing the Youden index), followed by Kaplan-Meier survival curves, log-rank tests, and univariate Cox regression to calculate hazard ratios for both PFS and OS at 4 years.

TL;DR: PARS operates in two stages: (1) PERCIST-based whole-body MFS detects all elevated-uptake regions above a blood-pool-derived SUVpeak threshold; (2) a CNN trained on 600+ lymphoma and lung cancer patients classifies each region as physiologic or suspicious. TMTVPARS is the sum of suspicious region volumes. No manual input is required. Statistical analysis used Bland-Altman, Spearman correlation, Kaplan-Meier, and Cox regression.
Pages 3-4
Region-Level Classification Accuracy: 85% Overall with 80% Sensitivity

Across the 280 patients, the MFS algorithm identified 6,737 ROIPARS sites with elevated FDG uptake. For comparison, the expert readers identified 7,996 ROIREF sites across the same cohort. Of the 6,737 automatically detected regions, the CNN classified 2,831 (42%) as suspicious uptake and 3,906 (58%) as physiologic. When compared against the expert reference, this yielded 2,399 true-positives (CNN correctly labeled as suspicious), 3,317 true-negatives (CNN correctly labeled as physiologic), 589 false-negatives (CNN labeled as physiologic but expert labeled as suspicious), and 432 false-positives (CNN labeled as suspicious but not matching expert regions). Overall region-level sensitivity was 80%, specificity was 88%, and accuracy was 85%.

Patient-level accuracy: At the per-subject level, mean classification accuracy was 87% (median 89%, interquartile range 81%-96%), indicating consistent performance across individual patients. On average, 24 regions per subject were detected by MFS, of which 20 were correctly classified (either as physiologic or suspicious) and 4 were misclassified. The median number of correctly classified ROIPARS sites was 17 (IQR 11-27), and the median number of incorrectly classified sites was 2 (IQR 1-5).

Spatial agreement between PARS and reference regions: The agreement between the set of PARS-classified suspicious regions and the expert-labeled ROIREF sites was further quantified using the Dice score, precision, and recall. The median Dice score was 0.73 (IQR 0.33-0.86), indicating moderate to good spatial overlap between the automated and manual region sets. Median recall (the fraction of reference lesion voxels captured in the suspicious PARS regions) was 0.62 (IQR 0.20-0.81), and median precision (the fraction of PARS suspicious voxels that overlapped with reference lesions) was 0.96 (IQR 0.86-0.99). The high precision alongside moderate recall indicates that PARS rarely incorrectly labeled physiologic tissue as tumor, but did miss some smaller or lower-uptake lesions that experts included in their reference segmentations.

Robustness across settings: Sensitivity analyses with different MFS threshold settings (2.5 SUV fixed threshold, and inclusion of sub-2 cm3 ROIs) showed that classification accuracy was not substantially impacted, demonstrating that the CNN performance was robust to the choice of initial region detection parameters.

TL;DR: The CNN achieved 80% sensitivity, 88% specificity, and 85% overall accuracy for classifying 6,737 detected regions across 280 patients. Per-subject accuracy was 87% (median 89%). Median Dice score between PARS suspicious regions and expert reference was 0.73. Precision was very high (median 0.96), indicating few false-positive lesion classifications, while recall of 0.62 reflects missed lower-uptake lesions.
Pages 3-5
TMTVPARS vs. Expert Reference: Strong Correlation with Systematic Underestimation

TMTVPARS values were substantially lower than TMTVREF across the cohort. The median TMTVPARS was 110 cm3 (IQR 33-281 cm3), compared to a median TMTVREF of 240 cm3 (IQR 80-529 cm3). The mean TMTVPARS was 235 cm3 (SD 348 cm3) versus 434 cm3 (SD 571 cm3) for TMTVREF, and the maximum observed TMTVPARS was 2,472 cm3 versus 3,833 cm3 for TMTVREF. Despite these absolute differences in volume, the rank correlation between the two methods was significant and strong, with a Spearman rank correlation coefficient of 0.76 (P less than 0.001), indicating that patients with high TMTV by one method tended to have high TMTV by the other.

Bland-Altman analysis: The Bland-Altman plot comparing the two methods showed wide limits of agreement, reflecting substantial case-by-case variability in the absolute volume differences. The Shapiro-Wilk test confirmed a non-normal distribution of differences (P less than 0.001), so the authors reported median bias and percentile-based limits of agreement rather than mean-based limits. The wide variability indicates that the two methods cannot be used interchangeably at the individual patient level, even though their prognostic performance is comparable at the population level.

Sources of underestimation: The authors identified several factors that likely contributed to TMTVPARS being lower than TMTVREF. First, the initial PERCIST-based threshold used in PARS (blood-pool-derived SUVpeak) is generally higher than the initial thresholds applied in the reference method, meaning some regions with modest uptake that the expert included in TMTVREF were never even detected by MFS. Second, the CNN systematically classified some lower-uptake lesions as physiologic that the expert labeled as suspicious. Third, the reference method included manual addition of missed lesions and drew prisms around suspicious regions using a 41% SUVmax threshold, which can contour larger volumes than the PARS approach. The lower TMTVPARS range simply shifted the optimal classification cutoffs downward without compromising prognostic discrimination.

No segmentation failures: Notably, no patients were excluded because of failure of the initial high-uptake region segmentation. The MFS algorithm identified at least one elevated-uptake region in all 280 subjects, confirming the robustness of the initial detection step even across diverse scanner hardware and reconstruction protocols from 124 centers.

TL;DR: TMTVPARS (median 110 cm3) was consistently lower than TMTVREF (median 240 cm3), but rank correlation was strong at r=0.76 (p less than 0.001). Bland-Altman analysis showed wide case-by-case limits of agreement. Underestimation stems from higher initial detection thresholds and CNN exclusion of lower-uptake lesions. MFS detected at least one high-uptake region in all 280 patients with zero segmentation failures.
Pages 4-5
TMTVPARS Predicts 4-Year PFS and OS Comparably to Expert Reference

The primary measure of clinical value in this study was whether automated TMTVPARS could stratify patient outcomes as effectively as expert-measured TMTVREF. For predicting 4-year PFS, the area under the ROC curve (AUC) was 0.63 for TMTVPARS and 0.69 for TMTVREF. The optimal TMTVPARS cutoff for PFS prediction was 171 cm3, while the TMTVREF cutoff was 242 cm3. These different thresholds reflect the systematic underestimation by PARS but achieve similar separation of risk groups. Using these cutoffs, 4-year PFS rates were 79% for low-TMTVPARS versus 54% for high-TMTVPARS, and 83% for low-TMTVREF versus 55% for high-TMTVREF. Log-rank tests confirmed significantly longer PFS for the low-TMTV group using both methods (P less than 0.001).

Overall survival results: For 4-year OS prediction, AUC was 0.65 for TMTVPARS and 0.68 for TMTVREF. Optimal OS cutoffs were 148 cm3 for TMTVPARS and 223 cm3 for TMTVREF. The 4-year OS rates were 90% for low-TMTVPARS versus 74% for high-TMTVPARS, and 93% for low-TMTVREF versus 74% for high-TMTVREF. Log-rank tests again showed significantly longer OS for low-TMTV patients by both methods (P less than 0.001).

Hazard ratios: Cox regression confirmed that both TMTV methods were independently predictive of outcomes. For PFS, the hazard ratio for high versus low TMTV was 2.3 (95% CI 1.5-3.6, P less than 0.001) using TMTVPARS and 2.6 (95% CI 1.6-4.1, P less than 0.001) using TMTVREF. For OS, the hazard ratio was 2.8 (95% CI 1.6-5.1, P less than 0.001) for TMTVPARS and 3.7 (95% CI 1.9-7.2, P less than 0.001) for TMTVREF. The slightly higher hazard ratios for TMTVREF likely reflect that the expert-measured volumes, though more time-intensive, captured the complete tumor burden with fewer missed lesions.

Clinical interpretation: Across all survival endpoints, TMTVPARS and TMTVREF performed comparably in terms of statistical significance and directionality of effect, confirming that the automated method retains the core prognostic information necessary for risk stratification in DLBCL. Sensitivity, specificity, negative predictive value, positive predictive value, and accuracy for predicting 4-year survival events were similar between the two methods at their respective optimal cutoffs.

TL;DR: TMTVPARS achieved AUC 0.63 for PFS and 0.65 for OS at 4 years, versus 0.69 and 0.68 for TMTVREF. Cox hazard ratios were 2.3 and 2.8 for TMTVPARS (PFS and OS) versus 2.6 and 3.7 for TMTVREF. Both methods showed P less than 0.001 on log-rank tests. Automated TMTV retains clinically meaningful prognostic discrimination comparable to expert manual measurement.
Pages 5-6
Methodological Constraints and Known Gaps

No gold standard for TMTV: A fundamental limitation of the study is that no universally accepted gold standard exists for TMTV calculation from 18F-FDG PET/CT. The expert reference method used in this study (Beth Israel Fiji plugin with manual correction) is itself just one of several approaches described in the literature. Because the figures of merit supporting PARS classification accuracy were calculated by comparison with this reference rather than an absolute ground truth, the results must be interpreted as agreement with one established method rather than as absolute accuracy. Methods may differ in which lesions they include (particularly low-uptake or small-volume sites), in contouring thresholds, and in the handling of diffuse bone marrow involvement.

Uniform lymphoma cohort: The REMARC trial enrolled a narrowly defined patient population: elderly (age 60-80), DLBCL-only, R-CHOP responders at study entry. The CNN was trained on a mixed cohort of lymphoma subtypes and lung cancer patients. Whether PARS performance generalizes to other lymphoma subtypes (follicular lymphoma, mantle cell lymphoma, T-cell lymphoma) or to younger DLBCL patients treated in different settings is not addressed by this data and may differ.

Systematic underestimation of TMTV: The consistently lower TMTVPARS values compared to TMTVREF mean that published TMTV cutoffs derived from the reference method cannot be directly applied when using PARS. Clinicians would need to recalibrate thresholds for each tool, complicating cross-study comparisons and the implementation of externally derived risk cutoffs in clinical practice.

Supervised use and motion artifacts: The authors acknowledge that PARS was originally designed for supervised use, where a reader can review and correct misclassified regions before finalizing the TMTV. Pitfalls in PET/CT image quality, including motion artifacts, patient movement during acquisition, and misregistration between PET and CT components, can degrade CNN classification outputs. In a fully automated deployment, these errors propagate directly into the TMTV estimate without opportunity for correction. Expert validation of PARS outputs is especially important when results are to be used for clinical risk stratification or treatment decisions.

TL;DR: Key limitations include the absence of a universal TMTV gold standard, cohort restriction to elderly R-CHOP-responding DLBCL patients, systematic underestimation requiring recalibrated cutoffs, and sensitivity to PET/CT image quality issues (motion, misregistration) that propagate unchecked in a fully automated pipeline. PARS was designed for supervised rather than fully automated clinical use.
Pages 6-7
Path to Routine TMTV Measurement and Risk-Adapted DLBCL Management

This study represents the first published demonstration that a fully automated AI method can generate TMTV values in a large DLBCL cohort with prognostic performance comparable to expert human measurement. The significance lies not only in the accuracy achieved but in the scalability: the pipeline requires no manual reader input, processes each scan in a fraction of the time needed for semiautomatic expert measurement, and was validated across 124 centers with heterogeneous scanner hardware and acquisition protocols. This robustness to real-world imaging variability is a critical prerequisite for any tool seeking clinical deployment across diverse practice settings.

Reducing observer variability: A persistent problem with current TMTV methods is interreader variability. Different nuclear medicine physicians applying the same software make different decisions about which regions to include, particularly for ambiguous foci near physiologic uptake sites (gut, bone marrow, joints). The PARS approach, by systematically applying a fixed decision boundary learned from training data, has the potential to reduce this variability and produce more reproducible TMTV estimates across centers and time points, which is important for multi-institutional trials and longitudinal monitoring.

Semi-supervised hybrid workflow: The authors propose an efficient middle-ground approach: use PARS to pre-classify all detected regions, then have an expert review only the small subset of ambiguous or potentially misclassified regions rather than performing the entire segmentation from scratch. On average, only 4 regions per patient were misclassified, meaning the expert would need to review and correct a modest number of sites rather than starting the classification of all 24 detected regions manually. This hybrid workflow could substantially reduce reading time while maintaining the quality control needed for clinical decision-making.

Broader TMTV standardization efforts: The authors situate this work within a larger field-wide initiative to standardize TMTV measurement for clinical use in lymphoma. Ongoing efforts include establishing guidelines for which anatomic regions should be included, defining consensus thresholding methods, and validating TMTV as a biomarker for treatment adaptation in prospective clinical trials. Automated tools like PARS will be essential components of this standardization effort, as they remove the subjectivity and resource burden that have prevented TMTV from becoming a routine clinical metric. Further development of deep-learning segmentation models trained specifically for DLBCL lesion detection, potentially incorporating transformer architectures and multi-scale attention mechanisms, may further improve sensitivity for smaller and lower-uptake lesions that PARS currently tends to miss.

TL;DR: PARS is the first fully automated AI system shown to generate prognostically valid TMTV in a large, multi-center DLBCL cohort. A hybrid human-AI workflow (AI pre-classifies all regions, expert corrects only the approximately 4 misclassified sites per patient) could cut reading time while preserving quality control. Standardizing TMTV measurement across institutions is a field-wide priority, and automated tools like PARS are a prerequisite for clinical adoption.