Soft tissue sarcomas (STS) present a distinctive challenge when oncologists try to evaluate whether chemotherapy is working. Unlike many solid tumors, STS do not reliably shrink in size during treatment even when they are biologically responding. Processes such as cystic degeneration, hyalinization, central necrosis, fibrosis, and intratumoral hemorrhage can all alter a tumor's internal composition without changing its outer dimensions measurably. This means that standard size-based response criteria, namely the World Health Organization criteria and the widely used RECIST (Response Evaluation Criteria In Solid Tumors), can misclassify a responding tumor as non-responsive simply because it did not shrink enough on imaging.
Choi criteria and their limitations: The Choi criteria, later extended as modified Choi criteria, were developed as an improvement by incorporating changes in CT attenuation or MRI signal intensity alongside size. These criteria were originally validated for gastrointestinal stromal tumors under imatinib therapy and demonstrated better correlation with pathologic response than RECIST in some settings. However, they were not designed specifically for STS, still rely heavily on size-based estimates, and have questionable applicability to morphologically complex subtypes like synovial sarcoma, where intratumoral signal heterogeneity is substantial and dynamic.
The case for radiomics: Radiomics is the systematic extraction of large numbers of quantitative imaging features from medical scans, transforming pixels into mineable multi-dimensional data. The defining appeal of radiomics over biopsy is that it can characterize the entire tumor volume non-invasively, capturing spatial heterogeneity that a needle biopsy samples only partially. In STS, radiomics has already been applied successfully to distinguish benign from malignant soft tissue masses, predict histologic grade, and estimate metastatic risk. The authors of this study hypothesized that measuring how radiomics features change between baseline and post-treatment MRI scans, a concept called delta-radiomics, might capture the biological changes associated with chemotherapy response that size criteria miss.
The clinical need is real: standard-of-care for newly diagnosed STS typically involves neoadjuvant anthracycline-based chemotherapy regimens that have demonstrated improved overall and metastasis-free survival in phase 3 trials, but no validated imaging biomarker yet exists that can reliably tell clinicians early in treatment which patients are responding and which are not. Collaborations between the FDA and the National Cancer Institute have explicitly called for quantitative imaging techniques to serve as surrogate biomarkers for treatment response in STS clinical trials.
This single-center retrospective study enrolled 44 patients (mean age 53.70 years, range 16 to 80 years; 43% male, 57% female) who received neoadjuvant chemotherapy (NAC) at the University of Southern California between January 2010 and January 2017. Patients were identified through chart review of cases presented at the institution's Orthopedic and Sarcoma Tumor Boards. Inclusion required both a baseline MRI prior to NAC initiation and a post-treatment MRI obtained at least 2 months after starting chemotherapy and before surgical resection. The institutional review board approved the study and waived informed consent given its retrospective nature.
Histologic composition: The cohort reflected real-world STS heterogeneity. The most common pathologic diagnosis was undifferentiated pleomorphic sarcoma (n = 17, 38.6%), followed by synovial sarcoma (n = 6, 13.6%), myxoid liposarcoma (n = 4, 9.1%), leiomyosarcoma (n = 4, 9.1%), myxofibrosarcoma (n = 2, 4.6%), extraskeletal Ewing sarcoma (n = 2, 4.6%), extraskeletal osteosarcoma (n = 2, 4.6%), and malignant peripheral nerve sheath tumor (n = 2, 4.6%), among other subtypes. Anatomically, lesions were concentrated in the thigh (n = 21, 47.2%), arm (n = 4, 9.1%), and pelvis or buttock (n = 4, 9.1%).
Multi-institutional MRI data: Two MRI scans per subject yielded 88 total studies. Notably, 37 scans were acquired at the authors' institution and 51 were acquired at 29 different outside facilities, creating a deliberately heterogeneous acquisition pool. The dataset encompassed 11 distinct MRI sequences, with T1, T2, and STIR sequences most commonly represented. This deliberate inclusion of multi-site data was a methodological choice intended to improve generalizability, reflecting real-world clinical practice where patients present with outside imaging rather than uniform in-house scans.
Tumor volumes were manually delineated as 3D regions of interest (ROIs) on one MRI sequence of interest per scan using Synapse 3D software (Fujifilm Medical Systems). These ROIs were then co-registered onto additional sequences using Statistical Parametric Mapping (SPM) software, allowing radiomics data to be extracted across multiple sequences simultaneously. From each 3D-ROI, the institutional radiomics pipeline extracted 1,708 features spanning 9 distinct texture families, including first-order statistics, gray-level co-occurrence matrices (GLCM), gray-level run-length matrices (GLRLM), and Laws Texture Energy (LTE)-derived metrics. The pipeline was rigorously benchmarked against an Image Biomarker Standardization Initiative (IBSI) phantom to ensure feature reproducibility.
Delta-radiomics calculation: For each patient, delta-radiomics features were computed as the arithmetic difference between post-NAC and pre-NAC feature values. This temporal subtraction approach is designed to capture the direction and magnitude of change in each texture characteristic across the treatment interval. The use of delta values, rather than static features from a single timepoint, is grounded in the principle that change in imaging texture may encode biological information about treatment-induced tumor remodeling that neither the baseline nor the follow-up scan conveys alone.
Machine learning classifiers: Two decision tree-based ensemble classifiers were trained: Random Forest (RF) and Real AdaBoost. RF was configured with 800 trees, leaf size of 16, maximal depth of 50, and the square root of the total variable count as the number of features sampled at each split. AdaBoost, being a more efficient algorithm, used only 25 trees with maximal depth of 3. Both models used Gini impurity as their loss function and incorporated prior correction to adjust for the imbalanced class distribution between responders and non-responders. Performance was evaluated using tenfold cross-validation, with AUC of the receiver operating characteristic (ROC) curve as the primary metric. Variable importance was ranked by out-of-bag Gini index, with a "cliff" in the Gini ranking used to define the threshold for top-performing features.
Statistical analyses: Univariate comparisons between NAC responders and non-responders used independent t-tests or Wilcoxon rank-sum tests depending on data normality. Multiple comparisons were corrected using the Benjamini-Hochberg procedure to control false discovery rate. All primary machine learning analyses were conducted in SAS Enterprise Miner 15.1 with High-Performance Procedures; all other statistical analyses used SAS v9.4.
The primary result of this study was negative. In the full, unfiltered machine learning analysis, both RF and AdaBoost failed to distinguish NAC responders from non-responders at a statistically meaningful level. RF produced an AUC of 0.40 (95% CI 0.22 to 0.58) and AdaBoost produced an AUC of 0.44 (95% CI 0.26 to 0.62). Both confidence intervals cross 0.50, meaning these models performed at or below chance, essentially equivalent to random classification. In the univariate analysis, only 4.74% of the 1,708 delta-radiomics variables (n = 265 features) reached statistical significance at p at or below 0.05, and only 1.34% (n = 75 features) reached significance at p at or below 0.01.
LTE features as a bright spot: Despite the overall negative result, the distribution of statistically significant features was not uniform across texture families. Laws Texture Energy (LTE)-derived metrics were markedly over-represented among the significant features: LTE accounted for 46.04% (n = 122) of all features reaching p at or below 0.05 significance, despite constituting only a portion of the total feature set. This finding is consistent with prior literature suggesting that spatial filtering techniques, which LTE represents, are particularly sensitive to voxel-to-voxel variation and intratumoral heterogeneity.
The filtered analysis experiment: As a proof-of-concept comparison exercise, the authors ran the machine learning procedure a second time after filtering the input features to only those that reached significance in the univariate analysis. When restricted to p at or below 0.05 features, RF achieved AUC 0.74 (95% CI 0.59 to 0.89) and AdaBoost achieved AUC 0.75 (95% CI 0.60 to 0.89). When restricted to p at or below 0.01 features, RF reached AUC 0.78 (95% CI 0.64 to 0.92) and AdaBoost reached AUC 0.82 (95% CI 0.70 to 0.95). These numbers are comparable to previously published positive studies in the field, but the authors explicitly flag this filtered approach as methodologically invalid, a point they develop at length in the discussion.
The authors contextualize their negative findings against a small but growing body of prior STS radiomics literature. The comparator studies all reported positive results but shared a common methodological feature: they applied feature selection or data filtering techniques before training their machine learning models, which the current study argues invalidates their performance metrics.
Crombé et al. (2019), NAC, n = 65, AUC 0.86: This study by Crombé and colleagues is the most directly comparable, as it also used an MRI-based delta-radiomics approach specifically for NAC response prediction in STS. They calculated the absolute change in 33 radiomics features in 65 STS patients following anthracycline-based NAC and trained RF, support vector machine, k-nearest neighbors, and logistic regression classifiers using 10-fold stratified cross-validation. Their highest AUC was 0.86. However, the current authors note that Crombé et al. constructed their models by first selecting one feature per category and then expanding via forward stepwise selection guided by univariate p-values, a procedure the current authors argue biases RF by pre-excluding features the algorithm is specifically designed to handle through its own internal weighting.
Peeken et al. (2021), neoadjuvant RT with or without NAC, n = 156, AUC 0.75: Peeken and colleagues used RF, LogitBoost, and elastic net regression with 3-fold nested cross-validation in 156 STS patients receiving neoadjuvant radiotherapy. Gao et al. (2020), neoadjuvant RT, n = 30, AUC 0.91: Gao and colleagues trained support vector machine and logistic regression classifiers using 5-fold cross-validation in a 30-patient cohort. Miao et al. (2022), neoadjuvant RT with or without TKI, n = 30, AUC 0.92: Miao and colleagues used logistic regression without any cross-validation procedure in a 30-patient cohort. All three of these studies employed either feature reduction techniques or lacked proper holdout validation, making their reported AUCs difficult to accept at face value.
The current authors demonstrate that when they apply the same kind of pre-selection filtering to their own dataset, they recover AUCs in the 0.74 to 0.82 range, directly matching those of the comparator studies. This strongly implies that the positive results in the prior literature are artifacts of methodological overfitting rather than genuinely predictive radiomic signal.
The most important methodological argument in this paper concerns information leakage: the inadvertent transfer of information from the test set into the model training process. In radiomics studies, information leakage commonly occurs when researchers perform univariate statistical testing across the entire dataset to identify significant features, and then use those pre-selected features as inputs for machine learning cross-validation. The critical problem is that the feature selection step already "saw" the test fold labels, so those test samples are no longer truly independent. RF and AdaBoost were both specifically designed to operate on high-dimensional feature sets without pre-selection because they have internal mechanisms (variable importance, bagging) that handle irrelevant features automatically. Using pre-selected features with these algorithms conflates the feature selection step with model training in a way that invalidates standard cross-validation estimates of generalization performance.
Publication bias in radiomics: The authors situate this methodological issue within a broader pattern of publication bias in the radiomics literature. A 2018 analysis by Buvat and colleagues found that only 6% of all PET radiomics studies in the published literature explicitly reported negative results. In a systematic review of 52 sarcoma-specific radiomics studies, Crombé and colleagues found that not a single study reported a specifically negative finding. This asymmetry is a well-recognized problem in medical research generally, but it is particularly acute in radiomics, where the combination of small sample sizes, high-dimensional feature spaces, and flexible machine learning pipelines creates enormous opportunity for inadvertent overfitting that produces optimistic performance metrics without genuine generalizability.
Why this study deliberately avoids filtering: The authors explicitly recommend against the routine use of feature reduction and data filtering in radiomics analyses and frame their unfiltered approach as the more rigorous, if less favorable, methodology. This stance is consistent with guidance from radiomics quality standards frameworks such as IBSI, which call for pre-registered analysis plans and independent external validation as prerequisites for clinical translation. The study's negative finding under these rigorous conditions is therefore not a methodological failure but rather an honest assessment of delta-radiomics performance under conditions that avoid bias.
Underpowered sample size: The authors openly acknowledge that 44 patients is likely insufficient to detect a significant radiomic signal with the statistical power needed for robust machine learning conclusions. A common benchmark in radiomics studies is 100 patients as the minimum threshold for reliable model training, particularly given the high dimensionality of radiomic feature spaces. STS is a rare malignancy, which makes assembling large single-institution cohorts inherently difficult. The authors note that achieving adequate sample sizes will likely require multi-institutional collaborations, but emphasize that such efforts must be paired with rigorous methodological standards, since simply pooling small biased analyses does not resolve the underlying overfitting problem.
Retrospective design and selection bias: Patients were identified through chart review of cases discussed at the institution's Orthopedic and Sarcoma Tumor Boards, which creates a risk for selection bias. Tumor board discussions tend to skew toward more complex, higher-risk, or diagnostically uncertain cases, meaning the enrolled cohort may not represent the full spectrum of STS patients who receive NAC in routine practice. Retrospective designs also lack the standardization of imaging protocols that prospective studies can enforce, contributing to the heterogeneity in acquisition parameters observed across the 29 contributing institutions.
Harmonization limitations: Despite using a multi-site dataset, the authors did not apply post-acquisition harmonization techniques such as ComBat, a statistical method designed to remove scanner and protocol batch effects from radiomic features. They explain that ComBat and similar approaches have meaningful limitations when applied to small samples with missing data and skewed distributions, all of which characterized their dataset. Without harmonization, scanner differences across 29 sites introduce signal noise that competes with any genuine biological signal in the delta-radiomics features. This is a fundamental tension in multicenter radiomics research: using multi-site data is necessary for generalizability, but harmonizing across sites remains technically challenging and may itself introduce artifacts in small cohorts.
The LTE signal: Despite the overall negative result, the authors point to LTE-derived features as a genuinely promising direction for future investigation. Laws Texture Energy is a spatial filtering technique based on convolution kernels derived from vector products of one-dimensional convolution masks, each representing a different texture property such as edge, spot, ripple, wave, or level. In this study's institutional pipeline, LTE-based metrics accounted for 1,472 of the 5,585 total features extracted across all 9 texture families. The fact that LTE features represented 46% of all statistically significant delta-radiomics variables, far above what would be expected by chance, suggests that these spatial domain descriptors are capturing something real about treatment-induced changes in intratumoral heterogeneity, even if that signal is not yet strong enough to drive reliable classifier performance at this sample size.
Histologic subtype stratification: The authors specifically call for future studies that correlate delta-radiomics changes with histologic subtype and pathologic findings of percent necrosis. Given that the enrolled cohort spans more than 10 STS subtypes with vastly different biologies, pooling them into a single classifier may obscure subtype-specific radiomics signals. A study powered to analyze undifferentiated pleomorphic sarcoma separately from synovial sarcoma, for instance, might reveal texture change patterns that are subtype-specific and therefore more predictive within each entity.
Chemotherapy regimen specificity: Another identified future direction is studies that stratify delta-radiomics changes by specific NAC regimen. Anthracycline-based, ifosfamide-based, and combination regimens produce different patterns of tumor remodeling at the histopathologic level. Texture features that reflect necrosis may behave differently under regimens with different mechanisms of cytotoxicity, and separating these regimen groups could reveal regime-specific radiomics signatures that a pooled analysis obscures.
Multi-institutional prospective design: The authors call for multi-institutional collaborations to achieve the sample sizes necessary for unfiltered machine learning analyses, particularly for rare STS subtypes. Future studies would benefit from prospective imaging protocol standardization, pre-registered analysis plans that commit to unfiltered machine learning approaches, and external validation cohorts that are completely held out from model development. Advances in post-acquisition harmonization methods, particularly those better suited to small and heterogeneous datasets than current ComBat implementations, will also be important enabling technologies for this research agenda.