Radiomics is the process of extracting large numbers of quantitative features from medical images obtained during routine clinical scans. Unlike visual reading by a radiologist, radiomics converts image data into hundreds or thousands of numerical descriptors of tumor texture, shape, and intensity that can then feed into machine learning models for diagnosis or outcome prediction. The analogy to other "omics" disciplines is intentional: just as genomics characterizes tumor DNA, radiomics aims to characterize the imaging phenotype of a tumor in a way that reflects its underlying biology.
The clinical translation gap: Despite a rapidly growing literature, radiomics remains almost entirely in the research phase. The reasons are methodological rather than conceptual. A 2021 systematic evaluation using the Radiomics Quality Score (RQS), developed by Lambin et al. in 2017, found that most published radiomic pipelines score poorly on comprehensiveness and reporting quality. Two distinct failure modes account for most of the gap. First, radiomic features are sensitive to how a region of interest (ROI) is drawn around the tumor: different readers or the same reader on different days can produce slightly different contours, and those contours can generate meaningfully different feature values. This is the reproducibility problem. Second, even a well-tuned model trained on one hospital's data may fail to generalize to another hospital's scanner settings, patient population, or histological mix. This is the validation problem.
Standardization initiatives: Two international efforts have tried to address these issues. The Image Biomarker Standardization Initiative (IBSI) established a stepwise consensus on how radiomic features should be calculated so that the same formula produces the same number regardless of the software used. The European Society of Radiology and European Society of Medical Imaging Informatics jointly published criteria for radiomic model development and a checklist called CLEAR (CheckList for EvaluAtion of Radiomics research) to guide both authors and peer reviewers.
Why sarcomas are a test case: Bone and soft-tissue sarcomas are rare and histologically heterogeneous cancers. Their low prevalence means individual institutions rarely accumulate enough cases for large-scale validation, which makes the reproducibility and validation challenges especially acute. A previous systematic review by the same author group covered studies published up to December 2020 and found that reproducibility analysis appeared in only 37% of papers and independent clinical validation in only 10%. The current review updates those figures for the period January 2021 through March 2023, covering the post-IBSI and post-CLEAR era.
The review was pre-registered on PROSPERO (CRD42023395542), following the convention for systematic reviews that aims to prevent post-hoc changes to eligibility criteria or outcome definitions. The search was conducted on two major biomedical databases: EMBASE (Elsevier) and PubMed (MEDLINE). The search window ran from 1 January 2021 through 31 March 2023, picking up exactly where the previous version of this review left off. The query used controlled vocabulary with medical subject headings and combined two conceptual domains: musculoskeletal sarcomas and radiomics, using the exact syntax: ("sarcoma"/exp OR "sarcoma") AND ("radiomics"/exp OR "radiomics" OR "texture"/exp OR "texture").
Reviewers and PRISMA compliance: The review involved three musculoskeletal radiologists with 3 to 5 years of experience in radiomics and sarcoma imaging (Gitto, Messina, Albano), who conducted literature search, study selection, and data extraction independently and in parallel. Disagreements were resolved by consensus involving a fourth radiologist with 8 years of AI and radiomics experience (Cuocolo). The Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) guidelines were followed, with the PRISMA checklist provided as supplementary material.
Inclusion criteria: Studies had to be original research published in peer-reviewed journals, focused on CT or MRI radiomics-based characterization of sarcomas in bone or soft tissues for either diagnostic or prognostic tasks, and required a statement confirming local ethics committee approval or adherence to institutional ethical standards.
Exclusion criteria: Papers were excluded if they focused on computer-assisted detection rather than mass characterization, if they dealt with retroperitoneal or visceral sarcomas, if they used animal or cadaveric models, if they were not in English, or if they had already been included in the 2020 version of this review. This last criterion specifically excluded papers published online in 2020 but appearing in print in a 2021 volume or issue.
Data extraction framework: Extracted variables were organized into four categories: baseline study characteristics (author, year, aim, tumor type, design, reference standard, imaging modality, database size, use of public data); segmentation and feature type (process, style, feature category); feature reproducibility strategies (method, statistical test, threshold); and model validation strategies (machine learning technique, internal test dataset, external test dataset). A pre-built spreadsheet with drop-down menus was used to standardize extraction and minimize categorical ambiguity.
After screening 201 papers identified through the search and applying eligibility criteria, 55 were included. Publication volume was nearly uniform across the review period: 24 studies (44%) in 2021, 23 (42%) in 2022, and 8 (14%) in the first quarter of 2023. The pace appears stable following the acceleration observed in the earlier review, where nearly half of all included studies were published in 2020 alone.
Study design and imaging modality: The overwhelming majority of included studies (54 out of 55, 98%) used a retrospective design. Only one study (2%) was prospective. MRI was the dominant imaging modality, used alone or in combination in 46 studies: exclusively MRI in 43 (78%), a combination of CT and MRI in 3 (6%), and CT alone in 9 (16%). The retrospective skew is largely driven by practical necessity, since bone and soft-tissue sarcomas are rare and radiology departments typically hold large archives of imaging data acquired during routine clinical management.
Tumor types: Studies were divided between bone sarcomas (n = 23) and soft-tissue sarcomas (n = 32). Within bone sarcomas, osteosarcoma and chondrosarcoma (including enchondroma vs. chondrosarcoma discrimination) dominated. Within soft-tissue sarcomas, the most common approach studied heterogeneous cohorts combining multiple histotypes. The median database size was 120 lesions (range 25 to 810). In three studies, multiple lesions per patient were counted separately, yielding 142, 128, and 161 lesions from 36, 125, and 160 patients, respectively. Public data were used in only 1 of 55 studies (2%), sourced from The Cancer Imaging Archive.
Study aims: Papers were roughly split between diagnostic and prognostic objectives. Diagnostic tasks included benign vs. malignant tumor discrimination (n = 20), grading (n = 8), histotype discrimination (n = 2), Ki-67 proliferation index prediction (n = 1), and marginal infiltration assessment (n = 1). Prognostic tasks covered survival prediction (n = 10), local or metastatic relapse prediction (n = 9), therapy response prediction to chemotherapy or radiotherapy (n = 11), treatment complication prediction (n = 1), and natural tumor evolution monitoring over time (n = 1). Several studies pursued two or three aims simultaneously.
Segmentation is the process of drawing a boundary around the tumor to define the region of interest from which radiomic features are computed. The way this is done matters enormously because every downstream feature value depends on which voxels are included in the mask. In the current review, manual segmentation was used in 48 studies (87%), semiautomatic segmentation in 5 (9%), a combination of manual and automatic in 1 (2%), and fully automatic segmentation in 1 (2%). One study used manual segmentation for handcrafted features while simultaneously extracting deep features from whole images without any explicit segmentation boundary, representing a hybrid approach.
Segmentation style (2D vs. 3D): Three-dimensional volumetric segmentation, where the entire tumor is delineated slice by slice, was used in 45 studies (82%). Two-dimensional segmentation on a single representative slice was used in 7 studies (13%) without multiple sampling and in 1 (2%) with multiple sampling. One study combined 3D and 2D approaches. In the 2D without multiple sampling group, most investigators chose the slice showing maximum tumor extension, although two studies used alternative criteria and one did not specify the selection rule.
Handcrafted vs. deep features: Handcrafted radiomic features are computed from predefined mathematical formulas applied to pixel intensity distributions within the ROI. These include first-order statistics (mean, standard deviation, entropy), shape metrics (volume, sphericity), and texture matrices (gray-level co-occurrence matrix, run-length matrix, gray-level zone size matrix). In contrast, deep features are learned representations extracted from intermediate layers of convolutional neural networks (CNNs), which capture spatial patterns without requiring explicit a priori formulas. Of the 55 included studies, 48 (87%) used only handcrafted features, 6 (11%) used both handcrafted and deep features, and 1 (2%) used deep features exclusively.
The small proportion using deep features reflects both the limited data available for training CNNs on rare sarcoma populations and the difficulty of interpreting what deep features represent biologically. Notably, the study using fully automatic segmentation combined it with deep feature extraction, suggesting that this approach benefits from end-to-end learning pipelines that do not require radiologist-drawn masks. As CNN architectures and publicly available sarcoma imaging datasets mature, a shift toward deep feature integration is likely.
Reproducibility analysis identifies which radiomic features produce consistent values when the same tumor is segmented multiple times under realistic conditions. Features that change substantially with minor segmentation differences are noisy rather than informative and should be excluded from model building. Among the 54 studies that used manual or semiautomatic segmentation (and therefore required human involvement where variability could enter), 32 (59%) included a formal reproducibility analysis. This represents a large increase from the previous version of this review, where only 18 of 49 eligible studies (37%) did so.
Strategies used: Inter- and intra-observer variability based on repeated segmentations was the approach in 30 of the 32 studies reporting reproducibility (55% of all eligible studies). In this approach, two or more radiologists delineate the same tumor independently (inter-observer) and/or the same radiologist delineates it again after a time gap (intra-observer), and the resulting feature values are compared. Two studies (4%) instead used geometrical perturbation of the ROI, applying small translations in multiple directions equal to 10% of the bounding box length, to simulate what different manual delineations might look like without requiring actual repeat segmentations. No study assessed the impact of image acquisition parameters (scanner model, field strength, reconstruction kernel) or post-processing techniques on feature reproducibility, which remains a significant methodological gap.
Statistical methods and thresholds: The intraclass correlation coefficient (ICC) was used in all 32 studies reporting reproducibility. ICC is a measure of agreement ranging from 0 to 1, where higher values indicate that a larger proportion of the total feature variance is attributable to differences between subjects (true signal) rather than differences between raters (noise). In the included studies, ICC threshold values for declaring a feature reproducible ranged between 0.7 and 0.9 depending on the study. These values are consistent with recent guideline recommendations for ICC-based reliability research. Additional methods used less commonly included the Bland-Altman method, Pearson correlation, and Spearman rank-order correlation. It should be noted that 7 studies validated segmentation accuracy using a second reader (reporting Dice similarity coefficients) without actually evaluating whether individual radiomic features were reproducible, and 3 additional studies reported only Dice scores without feature-level analysis.
The remaining 22 eligible studies (41%) performed no reproducibility analysis at all, meaning their feature sets may include noise-driven variables that appear predictive in training data simply because they encode segmentation idiosyncrasies rather than true tumor biology. This proportion is substantially lower than in the 2020 review baseline, but still represents a meaningful minority of the field.
Radiomic model validation occurs at two levels. The first is internal machine learning validation, which uses statistical resampling strategies during model development to estimate how well the model will perform on data it has not seen before. The second, and more clinically meaningful, level is clinical validation against a separate dataset that was held out entirely from model training. The distinction matters because resampling-based estimates can still be optimistic if the held-out samples come from the same institution, scanner, and patient mix as the training data.
Machine learning validation techniques: At least one resampling-based validation technique was used in 34 of 55 studies (62%), up from 25 of 49 (51%) in the 2020 review. K-fold cross-validation was the most commonly used approach, in which the dataset is divided into k equal partitions and the model is trained on k-1 folds and tested on the remaining fold, rotating until every sample has served as a test case. Less common techniques included bootstrapping (random sampling with replacement to create multiple training/test splits), leave-one-out cross-validation (the extreme case of k-fold where k equals the full sample size), and nested cross-validation (an outer loop for performance estimation and an inner loop for hyperparameter tuning, which prevents contamination between model selection and evaluation). One study used both k-fold and nested cross-validation in tandem.
Clinical validation against unseen data: Clinical validation was reported in 38 of 55 studies (69%), nearly double the 19 of 49 (39%) from the 2020 review. Internal test datasets, defined as held-out samples from the same institution separated by random split, temporal criteria, or different acquisition scanners, were used in 22 studies (40%). External test datasets from a different institution were used in 14 studies (25%). Two studies (4%) performed both internal and external validation. One notable case involved a multicenter study where patients were split randomly across institutions rather than by geography, which the authors categorized as an internal test on the grounds that training and test samples came from the same overall distribution.
The increase in external validation from 10% to 29% is the most encouraging finding in the review, as external validation is considered the gold standard for demonstrating that a radiomic model can generalize beyond its development context. Nevertheless, 31% of studies still reported no clinical validation of any kind, meaning roughly one in three papers provided no evidence that their model works on data it was not trained on.
The comparison between the 2020 and 2021-2023 review periods offers a concrete measure of how the field has evolved. Reproducibility reporting grew from 37% to 59% and clinical validation grew from 39% to 69%, with the external validation rate tripling from 10% to 29%. The authors interpret this as evidence that the radiomics community is actively responding to methodological criticism, with the IBSI, RQS, CLEAR checklist, and journal editorial policies collectively driving better practice.
Database size and public data: The median database size doubled from approximately 60 to 120 lesions between the two review periods, reflecting a trend toward larger cohorts. However, the use of public imaging data actually decreased, from 3 studies in the earlier review to only 1 in the current one (2%). This is concerning because public datasets are essential for independent model testing and comparison across research groups. The Cancer Imaging Archive, which was used in the one relevant study, hosts imaging collections for multiple cancer types, but dedicated sarcoma repositories remain sparse. The authors explicitly call for new publicly available sarcoma imaging databases.
Deep features and convolutional neural networks: Deep features, extracted from convolutional neural network layers rather than computed from predefined formulas, remain a minority approach (13% of studies). The authors note that fully automatic segmentation and deep feature extraction are natural companions, since end-to-end CNNs can learn both localization and characterization simultaneously without requiring manual ROI drawing. As large labeled sarcoma datasets become available through multicenter registries, deep learning methods are expected to grow. A 2022 study (Navarro et al.) using MRI of 306 soft-tissue sarcomas for grading with deep learning alone demonstrated the feasibility of this approach even in relatively small cohorts.
Remaining methodological gaps: Two reproducibility strategies that were present in the 2020 review baseline have disappeared entirely from the 2021-2023 literature: assessment of feature reproducibility under different acquisition techniques (scanner type, field strength, protocol) and under different post-processing approaches (reconstruction filter, slice thickness). These gaps matter because in multicenter studies, imaging protocol variation is a major source of non-biological radiomic variability. Addressing them prospectively requires controlled phantom studies or well-designed prospective multicenter protocols, neither of which appeared in the current review.
The authors identify several important limitations of their own systematic review. The most consequential is that no meta-analysis was performed. The reason is straightforward: the included studies are heterogeneous in their objectives, span multiple sarcoma histotypes, and use different performance metrics for model evaluation (AUC, sensitivity, specificity, accuracy, C-index, and others depending on task type). Even within a single task such as benign-malignant discrimination, the specific tumor types compared, the proportion of malignant cases, the imaging sequence used, and the feature selection strategy differ enough across studies to make pooled statistical estimates unreliable. This limits the review to descriptive synthesis rather than quantitative meta-analytic conclusions.
Reproducibility reporting gap: Even among studies that did assess feature reproducibility, most used ICC as a filter to remove non-reproducible features without actually reporting the ICC values or the proportion of features that survived the threshold. This prevents any cross-study quantification of how reproducible sarcoma radiomic feature sets tend to be. Future studies should report full ICC distributions or at least the fraction of features meeting each threshold.
Prospective design is nearly absent: With 98% of studies using a retrospective design, the evidence base for sarcoma radiomics is built almost entirely on data collected for clinical management rather than research purposes. Imaging protocols, scanner upgrades, and referral patterns introduce uncontrolled variability into retrospective cohorts. Prospective studies with standardized acquisition protocols and pre-specified endpoints are needed to generate the level of evidence required for regulatory approval or clinical guideline inclusion. The rarity of sarcomas makes prospective single-center studies impractical for achieving the sample sizes needed for robust external validation, pointing to the need for national or international multicenter registries.
Priorities for the next generation of studies: The authors identify three concrete priorities. First, larger multicenter investigations using geographically distinct training and validation cohorts, which would test whether radiomic models survive genuine institutional variation. Second, publication of imaging databases in open-access repositories such as The Cancer Imaging Archive, enabling independent replication and benchmarking. Third, increased use of prospective designs where possible, particularly for therapy response prediction tasks where pre-treatment imaging can be acquired under standardized conditions before treatment outcomes are known. The convergence of these three changes would accelerate the path from radiomic research to clinical decision-support tools in musculoskeletal oncology.