Correlation of Histopathology and Multi-Modal Magnetic Resonance Imaging in Childhood Osteosarcoma: Predicting Tumor Response to Chemotherapy

PLOS ONE 2022 AI 8 Explanations View Original
Original Paper (PDF)

Unable to display PDF. Download it here or view on PMC.

Plain-English Explanations
Pages 1-2
Why Predicting Chemotherapy Response in Osteosarcoma Is So Difficult

Osteosarcoma is the most common malignant bone cancer in children and adolescents, comprising roughly 60% of all bone sarcomas. It primarily strikes patients aged 10 to 20 years and carries a significant mortality burden for those whose tumors do not respond well to initial treatment. The standard-of-care protocol has remained essentially unchanged for three decades: patients receive pre-operative (neoadjuvant) chemotherapy, undergo surgical resection of the primary tumor, and then complete post-operative chemotherapy. The drugs used, including cisplatin, doxorubicin, and methotrexate, are cytotoxic agents that have shown efficacy but leave a substantial fraction of patients with tumors that fail to respond adequately.

The necrosis threshold: The single most important prognostic determinant for osteosarcoma is the degree of tumor necrosis observed in the surgically resected specimen after neoadjuvant chemotherapy. Pathologists grade response using histological examination: patients whose resected tumors show at least 90% necrosis are classified as "good responders" and have significantly better long-term survival than "poor responders" whose necrosis falls below that threshold. This histological grading has been the gold standard for treatment response evaluation for decades, but it has several critical limitations.

The problem with current evaluation: Standard tumor response metrics for other cancers, such as RECIST 1.1 (which measures dimensional shrinkage on serial imaging), are unreliable for osteosarcoma because the tumor's mineralized matrix does not shrink in response to cytotoxic agents even when chemotherapy is killing tumor cells. This means that MRI and CT cannot directly tell a clinician whether chemotherapy is working simply by tracking tumor size. Response is only confirmed at surgery, after 10 weeks of potentially ineffective treatment have already been administered. Beyond the delay, histological necrosis assessment also suffers from inter-observer variability, specimen loss during decalcification, and the fact that only a single representative section of a three-dimensional tumor mass is typically evaluated.

This study was designed to address that gap by developing a classification model that uses multi-modal MRI sequences, correlated directly to histopathology through image co-registration, to predict the degree of tumor necrosis before surgery and without the limitations inherent to tissue processing.

TL;DR: Osteosarcoma chemotherapy response is currently assessed only at surgery after 10 weeks of treatment, via histological necrosis grading (good response = at least 90% necrosis). RECIST criteria are unreliable for this tumor type. This study develops an MRI-based classification model to predict necrosis non-invasively before surgery, using multi-modal MRI correlated directly with histopathology.
Pages 3-5
Study Setup: A Single-Center Prospective Observational Cohort

The study was conducted at Children's Medical Center Dallas as a prospective, single-center observational study with institutional review board approval from the University of Texas Southwestern Medical Center. Patient enrollment ran from November 2016 to April 2019. Eligible patients were aged 10 to 30 years with newly diagnosed high-grade resectable osteosarcoma of an extremity (appendicular skeleton). Patients were excluded if they required sedation for MRI, had secondary bone sarcoma, had prior radiation to the affected field, or had a contraindication to contrast agent.

Enrollment and attrition: A total of 15 consecutive patients were enrolled (7 male, 8 female; ages 10 to 20 years, median 13 years). Tumor locations included the femur (5 cases), tibia (4 cases), humerus (3 cases), fibula (1 case), and radius (1 case). Two patients withdrew voluntarily, one became ineligible during the study, and two others were excluded because of inadequate pre-surgical imaging. This left n = 10 patients whose MRI data were used for classification model development. Of those 10, three had incomplete or invalid DCE-MRI sequences due to logistical or technical issues, reducing the sample further for analyses requiring dynamic contrast-enhanced sequences.

Chemotherapy protocol: All enrolled patients received the EURAMOS-I regimen, which includes cisplatin, doxorubicin, and methotrexate. The enrolled patients had advanced MRI sequences (diffusion-weighted imaging and DCE-MRI) added to their standard clinical conventional MRI acquisition at the pre-operative scan typically performed after week 10 of chemotherapy. An additional mid-course MRI (conventional plus DWI and DCE) was also performed after the fifth week of chemotherapy.

Power analysis context: A pre-study power analysis determined that 24 patients would be needed to provide 80% power to detect an odds ratio of at least 4.5 per standard deviation for a predictive feature. The target was not reached due to slow accrual from referral limitations and then disruption from the COVID-19 pandemic, which ultimately forced closure of enrollment. The final cohort of n = 10 is acknowledged as underpowered relative to the original design, and results are described as proof-of-concept rather than definitive validation.

TL;DR: 15 patients enrolled; 10 included in analysis (ages 10-20, median 13 years; femur most common site). All received EURAMOS-I chemotherapy (cisplatin, doxorubicin, methotrexate). Target sample of 24 was not met due to slow accrual and the pandemic. Study is powered as proof-of-concept. Three of the 10 patients had missing DCE-MRI data, reducing the DCE-inclusive analyses to n = 7.
Pages 5-7
Multi-Modal MRI Protocol: Conventional, Diffusion, and Dynamic Contrast Sequences

All MRI was performed on a 3-Tesla scanner (Magnetom Skyra, Siemens Healthcare) with images acquired in the coronal plane. The study used 11 distinct MR image types organized into three modality groups. Conventional MRI (CONV) consisted of three sequences: post-contrast T1-weighted spin-echo with fat saturation (PC), T1-weighted spin-echo (T1), and short inversion-time inversion recovery (STIR). Each of these had defined acquisition parameters (TR, TE, slice thickness, flip angle) selected for tissue contrast and anatomical detail.

Diffusion-weighted imaging: The readout-segmented echo-planar (RESOLVE) DWI sequence was acquired with three b-values: 0, 400, and 800 s/mm2. From these b-values, the apparent diffusion coefficient (ADC) was calculated for each pixel using the standard relationship S(b) = S0 exp(-b x ADC). ADC maps reflect the Brownian motion of water molecules in tissue, with necrotic regions expected to show elevated ADC values due to increased cellular permeability and reduced cellularity, while viable and more densely packed tumor tissue produces lower ADC values due to restricted diffusion.

Dynamic contrast-enhanced MRI: DCE-MRI was acquired using a T1-weighted volumetric interpolated breath-hold examination (VIBE) sequence. The gadolinium-based contrast agent (gadoterate meglumine, Dotarem) was injected intravenously at 2 mL/s immediately after the pre-contrast baseline frame was captured. DCE images were acquired continuously for three minutes at 19.5-second intervals. This temporal resolution allowed extraction of time-intensity curves for each image pixel. Four quantitative DCE-derived parametric maps were then computed: DCE-slope (steepest slope of the time-intensity curve), DCE-AUC (area under the time-intensity curve over the first 100 seconds), DCE-Ktrans (volume transfer constant between blood plasma and extravascular extracellular space, reflecting vascular permeability), and DCE-ve (relative extravascular extracellular space volume fraction). Three subtraction images (DCE-sub-0, DCE-sub-1, DCE-sub-2 at 0, 50, and 100 seconds after contrast arrival) were also generated. In total, the 11 MR image types comprised 3 conventional sequences, 1 ADC map, 4 DCE quantitative parameter maps, and 3 DCE subtraction images.

The two-compartment pharmacokinetic modeling for DCE-Ktrans and DCE-ve was performed in 3D Slicer using a defined arterial input function sampled from a feeding artery visible in each patient's DCE sequence, with contrast relaxivity set to 0.0039 L/mol/s for Gd-DOTA at 3T.

TL;DR: MRI acquired on a 3T Siemens scanner. Eleven image types total: 3 conventional sequences (PC, T1, STIR), 1 ADC map from DWI (b-values 0/400/800 s/mm2), 4 DCE quantitative parameter maps (slope, AUC, Ktrans, ve), and 3 DCE subtraction images. DCE-Ktrans and DCE-ve derived via two-compartment pharmacokinetic modeling in 3D Slicer.
Pages 7-10
Linking Tissue to Images: Deep Learning Pathology and Histologic-MRI Co-registration

Each surgically resected tumor specimen was sectioned into approximately 0.5-cm-thick slices in the coronal plane using a three-dimensional mold prepared from the mid-course (week 5) MRI to maintain anatomical orientation. Sections were fixed in 10% buffered formalin, decalcified with 10% hydrochloric acid, divided into 2 x 1.5 cm pieces, and processed into paraffin-embedded blocks. Individual 4-micrometer-thick histologic sections were cut, stained with hematoxylin and eosin, manually reviewed by a pathologist for necrosis estimation using the standard 90% threshold, and then digitally scanned using an Aperio AT Turbo scanner (Leica Biosystems) to produce whole slide images.

Deep learning tumor viability mapping: Rather than relying on manual pathologist delineation of necrotic regions within the digitized slides (which would be prohibitively time-intensive and subject to variability), a previously validated deep learning classifier was applied to each whole slide image. This model produced a tumor viability map for each slide, color-coding each region as either necrotic (red) or viable tumor (blue). Individual whole slide images were stitched together using dedicated image stitching software to reconstruct the full primary histologic section image and its corresponding tumor viability map.

Image co-registration: To correlate histopathology with MRI, the reconstructed histologic section image and its viability map had to be geometrically aligned to the corresponding MRI plane. This was accomplished using control point image mapping in MATLAB. A pathologist and a pediatric radiologist identified 8 to 12 control point pairs in each patient, selecting anatomical landmarks including the boundary of the tumor mass, the epiphysis, and the growth plate. A local weighted mean transformation was used to infer a geometric mapping from these control point positions, warping the histologic image onto the MR image coordinate system.

Multi-modal co-registration: The multi-modal MR images themselves were first aligned to a common coordinate system using the DICOM patient-based reference coordinate system and then underwent rigid body co-registration in 3D Slicer to correct for inter-scan patient motion. Signal intensities of conventional MR images were standardized using non-linear transformation to achieve consistency across patients. After co-registration, each pixel in the tumor area-of-interest on the MR image could be explicitly labeled as necrotic or viable based on the co-registered histologic viability map, creating the labeled training data for the subsequent classification model.

TL;DR: Histologic sections were digitized (Aperio AT Turbo) and analyzed by a pre-validated deep learning classifier to generate pixel-level tumor viability maps (necrotic vs. viable). These maps were then co-registered to MRI using 8-12 anatomical control points per patient and local weighted mean transformation in MATLAB, enabling direct pixel-level labeling of MR images based on histopathology.
Pages 10-13
Fuzzy C-Means Clustering and Weighted Majority Ruling for Necrosis Prediction

The core classification method consisted of two stages: feature extraction from MR images and multi-feature fuzzy clustering combined with a weighted majority rule. For feature extraction, a square pixel-centered window of side length W pixels was placed at each pixel within the co-registered tumor region. For each window position, 4 statistical parameters (original pixel value, mean, variance, skewness, kurtosis) and 19 Haralick texture features were computed. Haralick texture features are scalars derived from a gray level co-occurrence matrix (GLCM), which counts the co-occurrence of neighboring intensity values within a window and captures spatial relationships such as contrast, homogeneity, entropy, correlation, and energy. The GLCM was constructed as a direction-invariant matrix by summing eight directional GLCMs (left, right, up, down, and four diagonals).

Parameter optimization: The study systematically varied the window size W (3, 9, or 15 pixels) and the number of gray levels G (50, 100, or 200) used for GLCM construction to determine which settings best differentiated necrotic from viable tumor regions. One-way ANOVA was used to test which features were statistically significant (P less than 0.001 highly significant; 0.001 to 0.05 significant). A key finding was that window size had a meaningful impact on which features reached significance and the scale of spatial information captured, whereas varying the number of gray levels had little effect on significant feature count. Increasing W from 3 to 15 typically increased the number of significant features by 2 plus or minus 4, with notable exceptions for the PC, STIR, and DCE-sub-2 sequences.

Fuzzy c-means clustering: For each combination of window size, gray level count, and MR image type, fuzzy c-means (FCM) clustering was applied to the selected feature vectors. FCM partitions pixels into c = 2 overlapping clusters (necrotic and viable), assigning each pixel a continuous membership value between 0 and 1 indicating its degree of belonging to each cluster. The algorithm minimizes an objective function iteratively, updating cluster centers and membership values until convergence (sensitivity threshold epsilon = 10^-5). The use of fuzzy membership rather than hard binary assignment is well-suited to heterogeneous osteosarcoma tissue where the boundary between necrosis and viable tumor is not always sharp.

Weighted majority ruling: For a given MRI subset (a defined combination of MR image types), the individual FCM fuzzy segmentation maps were combined using a weighted majority rule. For each pixel, if the weighted sum of its membership values to necrosis across all maps in the subset exceeded half the total number of maps, it was classified as necrotic in the final binary segmentation. Optimal weights were determined by minimizing either the mean absolute error (min-avg) or the maximum absolute error (min-max) between MRI-predicted necrosis percentage and pathologist-estimated necrosis percentage, using a brute-force discrete search across all valid weight permutations. Eight MRI subsets were evaluated, ranging from CONV alone (3 images) to the full combination of all 11 images.

TL;DR: Each MR image pixel is described by 24 features (4 statistical plus 19 Haralick texture), computed from a sliding window (W = 3, 9, or 15 pixels; G = 50, 100, or 200 gray levels). Fuzzy c-means clustering (c = 2 clusters, fuzziness exponent m = 2) assigns pixel-level necrosis probability. A weighted majority rule combines maps from different MRI modalities, with weights optimized by brute-force minimization of mean or maximum absolute error vs. pathologist necrosis estimates.
Pages 13-16
Key Findings: Accuracy by MRI Modality and the Role of DCE-MRI

The primary outcome measure was the absolute error between MRI-predicted tumor necrosis percentage and pathologist-estimated necrosis percentage. The results varied substantially depending on which MRI sequences were included in the classification model. The study's most important finding was that conventional MRI alone achieved an average accuracy above 90% in differentiating necrosis from viable tumor. Specifically, using the CONV subset (PC, T1, and STIR sequences combined), the absolute error in necrosis estimation was 8.2 plus or minus 1.7% (n = 10). Diffusion-weighted imaging alone (ADC) produced a similarly modest absolute error, and critically, adding ADC to conventional MRI (CONV + DW) did not statistically improve accuracy over CONV alone.

The impact of DCE-MRI: Adding the four DCE quantitative parametric maps (DCE-q: slope, AUC, Ktrans, ve) to conventional MRI produced a statistically significant improvement in accuracy (P at or below 0.03 for CONV vs. CONV + DCE-q and related comparisons). The CONV + DCE-q subset reduced the mean absolute error meaningfully, and the best-performing subsets were CONV + DCE-q, CONV + DW + DCE-q, and CONV + DW + DCE-q + DCE-s. For the CONV + DW + DCE-q combination, the absolute error was 4.1 plus or minus 0.8% (n = 7), roughly halving the error seen with conventional MRI alone. Adding the DCE subtraction images (DCE-s) to CONV + DW + DCE-q did not provide any additional statistically significant improvement, suggesting that DCE-slope and DCE-AUC captured the essential spatial contrast enhancement information already present in the raw subtraction frames.

Spatial similarity of necrosis maps: In addition to the percentage-based absolute error, the spatial correspondence between MRI-predicted and histology-derived necrosis regions was evaluated using Dice and Szymkiewicz-Simpson (overlap) coefficients. For CONV + DW + DCE-q (min-avg optimization), the Dice coefficient was 0.70 plus or minus 0.03 and the overlap coefficient was 0.94 plus or minus 0.01 (n = 7). The consistently high overlap coefficient (close to 1.0 across MRI subsets) indicated that histologically confirmed necrotic pixels were largely contained within the MRI-predicted necrotic region. The lower Dice coefficient (0.61-0.70) reflected the fact that the MRI-based model identified a larger necrotic region than the histologic viability map, partly because void areas from liquefied tissue lost during specimen processing were counted as non-necrotic in the pathology analysis but were correctly predicted as necrotic by MRI.

Optimization method comparison: Between the two weight optimization approaches, min-avg consistently produced lower mean absolute error but higher variability (larger standard error of the mean) compared to min-max, which sacrificed mean accuracy for lower variance. Neither method was statistically superior to the other for any given MRI subset, suggesting both are viable depending on whether minimizing average or worst-case error is the clinical priority.

TL;DR: Conventional MRI alone achieved absolute error of 8.2 +/- 1.7% (n = 10; accuracy above 90%). Adding DCE-q (slope, AUC, Ktrans, ve) significantly improved accuracy to 4.1 +/- 0.8% (n = 7) for CONV + DW + DCE-q. Adding ADC alone to CONV gave no benefit. Best spatial overlap coefficient: 0.94 +/- 0.01. Dice coefficient: 0.70 +/- 0.03 for CONV + DW + DCE-q.
Pages 16-18
Whole-Tumor Volume Analysis vs. Single-Section Histopathology

Standard histopathological necrosis grading examines a single primary section of tumor, selected by the pathologist as the most representative coronal plane. This is an inherent limitation: a three-dimensional tumor mass with heterogeneous necrosis distribution may be poorly characterized by any single cross-section. In this study, once the classification model was established using the plane-of-interest (POI) analysis, it was extended to predict tumor necrosis across the entire tumor volume-of-interest (VOI), segmented by the pediatric radiologist on the post-contrast T1-weighted sequence.

VOI versus AOI findings: When necrosis estimates from the entire tumor VOI (NVOI) were compared to those from the single best-matching MRI plane (NAOI) and to the pathologist's histological estimate (Nhist), a consistent pattern emerged. Volume-based necrosis estimations were generally lower than single-plane estimates. While this difference did not reach statistical significance across the board, the directional trend was clear and consistent. This is clinically meaningful: it implies that a single histologic or MR section may systematically overestimate the proportion of the tumor that is necrotic, because pathologists typically select the section that appears most necrotic. A volumetric assessment, which integrates spatial information from the entire tumor, would be expected to give a more accurate and less biased representation of actual tumor response.

Three-dimensional viability maps: The study generated three-dimensional tumor viability maps showing the spatial distribution of necrotic and viable regions across all coronal slices of the tumor VOI. These maps revealed, for individual patients, that the distribution of necrosis in the full three-dimensional volume could differ substantially from what was apparent in a single representative slice. For good responders (Nhist at or above 90%), the three-dimensional maps showed predominantly red (necrotic) voxels, while poor responders showed larger regions of blue (viable tumor) extending through the volume.

The ability to generate a volumetric necrosis map before surgery, using only MRI, has potential implications for surgical planning (by identifying which parts of the tumor retain viable cells), for prognostication (by providing a more complete picture of tumor biology), and for real-time treatment adaptation (by potentially detecting early response or resistance after 5 weeks rather than waiting for surgery at 10 weeks).

TL;DR: Volume-based necrosis estimates (from the full tumor VOI) were consistently lower than single-plane estimates from the matched histologic section or the MRI plane-of-interest, though the difference was not statistically significant across all comparisons. This directional pattern suggests that single-section histopathology may overestimate tumor necrosis. The model can generate full 3D tumor viability maps, which could inform surgical planning and earlier response assessment.
Pages 18-21
Limitations, Methodological Caveats, and the Path to Clinical Translation

Small sample size: The most significant limitation of this study is the small cohort of n = 10 patients available for model development, reduced further to n = 7 for analyses requiring complete DCE-MRI data. Osteosarcoma is a rare disease, and high-grade appendicular presentations meeting the strict eligibility criteria are rarer still. The authors note that almost all patients enrolled were good responders with tumor necrosis at or above 90%, which may bias the weight optimization procedure toward identifying necrosis rather than distinguishing good from poor responders. Only with a larger, more balanced dataset (including more poor responders) could the model's predictive performance for the clinically critical good-vs-poor-response classification be robustly estimated.

Tissue processing artifacts: A recurring challenge in correlating MRI with histopathology for osteosarcoma is that liquefied necrotic tissue is frequently lost during specimen processing (fixation, decalcification, and block preparation). This tissue shrinkage and loss means the histologic viability map produced by the deep learning classifier is inherently incomplete, underrepresenting the true extent of necrosis. This artifact contributes to the comparatively high absolute errors associated with the histologic deep learning reference method itself (visible in the study's figures), even though the classifier correctly identified the necrotic tissue that was present in the sections. The authors argue this artifact actually underscores the advantage of MRI-based necrosis quantification, which is not subject to tissue loss.

Co-registration imperfections: Achieving a perfect geometric alignment between a histologic section and an MRI plane is technically challenging due to tissue deformation during processing, the two-dimensional nature of a histologic section cut from a three-dimensional specimen, and the inherent differences in image resolution between digitized pathology slides and MRI. The use of 8-12 control points with local weighted mean transformation provided reasonable alignment, as verified by a pathologist and radiologist, but residual misalignment contributes to error in the feature labeling step.

Future directions: The authors identify several concrete paths forward. The current fuzzy c-means and weighted majority rule framework is relatively simple, chosen partly because it is interpretable and feasible with n = 10 patients. As more imaging data becomes available through expanded enrollment or multi-center collaboration, more sophisticated approaches such as convolutional neural networks or other deep learning architectures could be applied to the raw MR images directly, bypassing the need for explicit feature engineering. An alternative weight optimization objective combining both mean and maximum absolute error minimization could offer a better balance than either criterion alone. The authors also propose investigating volumetric histological grading (from multiple sections rather than one) as a more accurate reference standard against which to validate the volumetric MRI-based necrosis estimates.

TL;DR: Key limitations: n = 10 (underpowered), most patients were good responders (biasing toward necrosis identification), tissue loss during decalcification creates incomplete histologic reference maps, and co-registration imperfections introduce labeling errors. Future work will apply deep learning architectures as sample sizes grow, explore multi-objective weight optimization, and validate volumetric MRI necrosis estimates against multi-section histological grading.