Machine Learning-Supported Diagnosis of Small Blue Round Cell Sarcomas Using Targeted RNA Sequencing

Journal of Molecular Diagnostics 2024 AI 8 Explanations View Original
Original Paper (PDF)

Unable to display PDF. Download it here or view on PMC.

Plain-English Explanations
Pages 1-2
The Diagnostic Challenge of Small Blue Round Cell Sarcomas

Small blue round cell sarcomas (SBRCSs) are a heterogeneous group of mesenchymal tumors that predominantly affect children and young adults. Their name derives from the appearance of their cells under a microscope: small, round to ovoid cells with blue-staining nuclei in hematoxylin and eosin staining. The clinical problem is that multiple distinct tumor subtypes share this morphologic appearance, yet they carry profoundly different prognoses and require different treatments. Distinguishing them from one another is one of the most challenging tasks in soft-tissue sarcoma pathology.

The four main SBRCS subtypes: Ewing sarcoma (ES) is the most common type, defined by gene fusions between EWSR1 and transcription factors of the ETS family, predominantly EWSR1::FLI1 (85% of cases) and EWSR1::ERG (10%). The remaining cases belong to a group called Ewing-like sarcomas, which resemble classic ES morphologically but lack the EWSR1::ETS molecular hallmark. The four recognized Ewing-like subtypes are CIC-rearranged sarcomas, BCOR-rearranged sarcomas, sarcomas with EWSR1::non-ETS fusions, and round cell sarcomas with other gene fusions.

Why subtypes matter clinically: Although these tumors look similar under the microscope, their clinical behavior is very different. CIC-rearranged sarcomas, for example, carry a significantly worse prognosis than classic Ewing sarcoma, with poorer responses to standard Ewing chemotherapy regimens. BCOR-rearranged sarcomas behave differently still. Correctly identifying the subtype guides treatment selection, inclusion in clinical trials, and prognostic counseling for patients and families.

The molecular diagnostic bottleneck: The current gold standard for SBRCS classification relies on detection of the specific chromosomal rearrangements (gene fusions) that define each subtype. However, a major practical problem exists: the most common CIC rearrangement, CIC::DUX4, involves a translocation between chromosome 19 and chromosome 4q35, a region containing repetitive DNA sequences. This repetitive structure makes the CIC::DUX4 fusion notoriously difficult to detect by standard next-generation sequencing (NGS) approaches and FISH assays. Up to 50% of CIC::DUX4 cases can be missed by routine RNA fusion panels. This diagnostic gap is what the study aims to close.

TL;DR: SBRCSs are morphologically similar but molecularly distinct sarcoma subtypes with very different prognoses. CIC-rearranged sarcomas are particularly hard to diagnose because the CIC::DUX4 fusion is frequently missed by standard sequencing and FISH, motivating a machine learning approach based on gene expression patterns rather than direct fusion detection.
Pages 2-3
Study Design: Two Cohorts and a Routine Sequencing Panel

The study was conducted at the Institute of Pathology in Erlangen, Germany, and used tumor samples analyzed during routine clinical diagnostics. The experimental framework relied on the Illumina TruSight RNA Fusion panel, a targeted RNA hybrid capture sequencing platform designed for routine use in oncology. Rather than developing a novel bespoke assay, the researchers asked whether the expression data already generated as a byproduct of routine fusion testing could support a higher-level diagnostic classification.

Curated cohort (training and development): 69 soft-tissue tumors (STTs) with confirmed diagnoses reviewed by an expert STT pathologist (A. Agaimy) were assembled. This group included 31 SBRCSs and 38 other STT types for context. The 31 SBRCSs comprised 13 classic Ewing sarcomas, 3 BCOR-rearranged sarcomas (confirmed by BCOR immunostaining), and 15 cases with morphologic features of CIC-rearranged sarcoma. Of those 15 CIC-morphology cases, RNA sequencing identified gene fusions in only 10, confirming the CIC::DUX4 detection problem: 5 cases had the expected CIC-rearranged morphology but no detectable fusion.

Test cohort (validation): The remaining 1,335 routine diagnostic cases from the same institution formed the test cohort, covering primarily STTs, salivary gland carcinomas, and kidney tumors. These cases were analyzed with the same Illumina TruSight Panel but without the benefit of expert STT pathologist review at the time of initial diagnosis. The classifier was then applied prospectively to these 1,335 cases to identify any unrecognized CIC-rearranged sarcoma candidates.

RNA extraction and sequencing: Tumor RNA was isolated from microdissected formalin-fixed, paraffin-embedded (FFPE) tissues (approximately five sections, 6 to 8 micrometers thick) using the Qiagen RNeasy FFPE kit. Only samples with more than 30% of RNA fragments exceeding 200 nucleotides passed quality control. Libraries were prepared with the Illumina TruSight RNA Fusion Panel v2.0 and sequenced on a NextSeq 500 instrument with 75-base paired-end reads to a minimum of 10 million reads per sample. A subset of samples was also run on the Archer FusionPlex Sarcoma Kit using 250 ng total RNA, sequenced on a MiSeq with 151-base paired-end reads to a minimum of 1.5 million reads per sample.

TL;DR: The study used 69 expert-reviewed STTs as a development cohort and 1,335 routine diagnostic cases as a test cohort, both sequenced with the Illumina TruSight RNA Fusion Panel from FFPE tissue. Key design insight: the classifier reuses expression data already generated during standard fusion testing, adding diagnostic value at no additional cost.
Pages 3-4
What Targeted RNA Sequencing Detected (and Missed)

In the 38 non-SBRCS soft-tissue tumors, targeted RNA sequencing detected the expected gene fusion in every single case, confirming the panel's reliability across diverse STT subtypes. This included: SS18::SSX1 in 8 of 12 synovial sarcomas, SS18::SSX2 in 4 of 12 synovial sarcomas, NAB2::STAT6 in all 10 solitary fibrous tumors, FUS::DDIT3 in all 6 myxoid liposarcomas, and COL1A1::PDGFB in all 6 dermatofibrosarcoma protuberans (DFSP) cases. The 4 DFSP cases with fibrosarcomatous transformation also showed the same COL1A1::PDGFB fusion. The panel's sensitivity was thus 100% across these non-SBRCS entities.

SBRCS fusion detection: the CIC problem becomes concrete: Among the 31 SBRCS samples, RNA sequencing identified the expected gene fusions in all 13 classic Ewing sarcomas (10 with EWSR1::FLI1, and one each with EWSR1::ERG, EWSR1::FEV, and FUS::ERG). All 3 BCOR-rearranged sarcomas were identified by their BCOR::CCNB3 fusion. However, among the 15 cases with CIC-rearranged sarcoma morphology, RNA sequencing detected a gene fusion in only 10, leaving 5 with no identified molecular marker. These 5 unresolved cases were designated as "CIC-rearranged-like" sarcomas.

FISH confirmation: To validate the presence or absence of CIC rearrangements, FISH using a CIC gene locus dual color break-apart probe was performed. The FISH results confirmed that all 10 CIC-fusion-positive cases had rearrangements, while the 5 "CIC-rearranged-like" cases had negative FISH results. This pattern is consistent with published literature showing that CIC::DUX4 translocations are detectable by FISH in only about 60 to 70% of cases, and that some cases lack detectable CIC rearrangements despite identical histomorphology and gene expression profiles.

Implications for the expression-based approach: The failure of fusion detection in the "CIC-rearranged-like" cases actually sets up the central argument for the machine learning classifier: if the gene expression profile of these 5 unresolved cases resembles the 10 confirmed CIC-rearranged cases, then expression profiling can serve as a surrogate diagnostic marker even when direct fusion detection fails.

TL;DR: RNA sequencing achieved 100% fusion detection in non-SBRCS tumors, 100% in Ewing sarcoma, and 100% in BCOR-rearranged sarcoma, but only 67% (10 of 15) in CIC-morphology cases. The 5 undetected cases are likely true CIC-rearranged sarcomas based on their expression profiles.
Pages 4-5
Distinct Gene Expression Landscapes Across SBRCS Subtypes

Before training any classifier, the researchers explored whether the expression data from the TruSight panel naturally separated the SBRCS subtypes without supervised guidance. Using variance-stabilized transformed (VST) counts from the DESeq2 package in R/Bioconductor, they identified the 100 genes with the highest expression variability across all 69 curated cohort samples. Principal component analysis (PCA) of these 100 genes revealed that tumor entities clustered strongly by diagnosis, with each STT subtype occupying a distinct region in principal component space.

CIC-rearranged and CIC-rearranged-like cases cluster together: A particularly important finding was that the 5 CIC-rearranged-like cases, which lacked any detectable fusion, clustered directly alongside the 10 confirmed CIC-rearranged sarcomas in PCA space. This was not guaranteed: if the CIC-rearranged-like cases were truly molecularly distinct entities, they would be expected to cluster separately. Their co-localization with the confirmed CIC-rearranged cases supports the interpretation that they are biologically part of the same group, with a technically undetectable CIC::DUX4 fusion.

Differential expression confirms the separation: Formal differential gene expression analysis using DESeq2 confirmed quantitatively what the PCA showed visually. Between CIC-rearranged sarcomas and BCOR-rearranged sarcomas, 68 genes were differentially expressed at an adjusted P value below 0.01. Between CIC-rearranged sarcomas and Ewing sarcomas, 137 genes were differentially expressed at the same threshold. In both comparisons, ETV1, ETV4, and WT1 were among the most strongly upregulated genes in CIC-rearranged sarcomas, consistent with prior literature identifying ETV4 as a potential immunohistochemical marker for this subtype.

Gene-level expression dot plots: Individual gene expression plots confirmed highly specific upregulation patterns: CCNB3 was selectively expressed in BCOR-rearranged sarcomas, PDGFB was selectively elevated in DFSP, SSX1 distinguished synovial sarcomas, and ETV1 and ETV4 stood out in CIC-rearranged and CIC-rearranged-like tumors. WT1 was also elevated in the CIC group but showed broader expression, foreshadowing its limited specificity when used alone.

TL;DR: PCA of the top 100 variable genes separated all SBRCS subtypes cleanly without supervision. The 5 CIC-rearranged-like cases without detectable fusions clustered with the confirmed CIC-rearranged cases. DESeq2 identified 68 differentially expressed genes between CIC-rearranged and BCOR-rearranged sarcomas, and 137 between CIC-rearranged and Ewing sarcomas (adjusted P below 0.01), with ETV4 consistently prominent.
Pages 5-7
Building and Validating the Random Forest Classifier

The core machine learning contribution of the paper is a random forest (RF) classifier trained to predict the probability that a given tumor is CIC-rearranged. Random forests are ensemble learning methods that build many decision trees on random subsets of training data and features, then combine their predictions by majority vote. This approach is well suited to high-dimensional genomics data with relatively small sample sizes because it is robust to overfitting and provides interpretable feature importance metrics.

Training set construction: The classifier was trained exclusively on CIC-rearranged and Ewing sarcoma cases from the curated cohort, because these represent the clinically most critical differential diagnosis. Raw gene counts were transformed into VST counts, and the top 100 variable genes across these cases were identified. The VST counts of those 100 genes were used as input features. The random forest was trained with 1,000 trees (ntree = 1000). Leave-one-out cross-validation within the training set (out-of-bag, OOB, estimation) was used to assess classifier performance without requiring a separate held-out set, given the limited training sample size.

Training performance: In the OOB predictions on the training set, all CIC-rearranged cases received very high predicted probabilities of being CIC-rearranged, clearly separated from Ewing sarcoma cases, which received probabilities near zero. The classifier was then applied to the curated cohort samples not included in training, confirming good separation. Critically, the 5 CIC-rearranged-like cases (without detected fusions) received high predicted probabilities, reinforcing their classification as likely CIC-rearranged sarcomas.

Variable importance: The RF variable importance plot identified ETV4 as the single most important predictor, followed by a set of additional genes. The use of a 100-gene panel, rather than ETV4 alone, proved essential: the RF achieved markedly better discrimination than ETV4 expression alone because many cases with high ETV4 expression due to other reasons (roughly 22 additional cases in the test cohort exceeded the 97th percentile ETV4 threshold) were correctly classified as non-CIC-rearranged by the multigene model. The RF essentially learned that CIC-rearranged sarcomas have a coordinated transcriptional program, not just a single elevated marker.

TL;DR: A random forest trained on VST counts of the top 100 variable genes from CIC-rearranged and Ewing sarcoma cases cleanly separated these subtypes using OOB cross-validation. The 5 fusion-negative CIC-rearranged-like cases received high classifier probabilities, supporting their re-classification. ETV4 was the top feature, but the multigene model substantially outperformed ETV4 alone.
Pages 7-9
Prospective Testing Across 1,335 Routine Diagnostic Cases

The most clinically important part of the study is the application of the trained classifier to 1,335 routine diagnostic tumor samples processed at the same institution. These samples were not hand-picked for SBRCS content; they represent a realistic distribution of tumors encountered in day-to-day pathology practice, including STTs, salivary gland carcinomas, and kidney tumors. This large-scale test reflects how the tool would actually perform if deployed in a clinical laboratory.

Distribution of predicted probabilities: For the vast majority of the 1,335 test cases, the classifier assigned a predicted probability of being CIC-rearranged very close to zero, reflecting that most tumors in an unselected cohort are not CIC-rearranged sarcomas. The histogram of predicted probabilities showed a strongly right-skewed distribution with a large spike near zero and a small tail of high-probability cases. A probability threshold of 0.75 was chosen by visual inspection of this distribution to flag 15 samples as candidate CIC-rearranged cases.

The 15 candidate cases: These 15 high-probability cases are presented in Table 1 of the paper and carry a range of predicted probabilities from 0.94 to 0.99, with two additional cases receiving probabilities of 0.94. Initial pathology diagnoses for these cases included "unclassified highly malignant epithelioid neoplasm," "Ewing-like sarcoma, most likely CIC-rearranged," "undifferentiated round cell sarcoma, NOS," "unclassified highly malignant round and spindled cell sarcoma," and other non-specific descriptions, reflecting the diagnostic difficulty that motivated the study. One additional borderline case, S_09, received a probability of 0.72 and was flagged separately as meriting follow-up investigation.

Unsupervised clustering confirms the candidates: When the expression profiles of the 15 candidate test cases were analyzed alongside the training cohort using unsupervised clustering of the top 20 RF predictor genes, the candidates clustered directly with the confirmed CIC-rearranged training cases, providing independent confirmation of the classifier's assignments. All 15 candidates showed very high ETV4 expression. However, they did not have the highest WT1 values overall, illustrating that ETV4 was the dominant driver of classification while WT1 added secondary resolution to distinguish true CIC-rearranged cases from non-CIC cases with spuriously high ETV4.

TL;DR: Applied to 1,335 routine cases, the classifier identified 15 high-confidence candidate CIC-rearranged sarcomas (predicted probability above 0.75) that had been ambiguously classified on initial pathology review. Unsupervised re-clustering of the 20 top predictor genes confirmed these candidates clustered with known CIC-rearranged sarcomas, validating the classifier's prospective utility.
Pages 9-10
The Advantage of Multigene Profiling Over Single Marker Testing

One of the paper's most practically useful findings concerns the limitations of single-gene marker approaches. Published studies have proposed ETV4 immunohistochemistry or mRNA expression as a convenient surrogate for CIC rearrangement, since ETV4 is markedly upregulated in this subtype. This study provides quantitative evidence for why single-marker approaches are insufficient when used in isolation.

The ETV4 threshold problem: The researchers defined a high-ETV4-expression threshold as VST counts above 11, which corresponded to approximately the 97th percentile of the full 1,335-case test cohort. This threshold correctly captured all 15 confirmed candidate CIC-rearranged cases. However, it also captured 22 additional cases whose ETV4 expression exceeded the cutoff but whose overall expression profiles did not match CIC-rearranged sarcomas. These 22 "false positive" cases received markedly lower RF probabilities, correctly indicating they were not CIC-rearranged sarcomas despite high ETV4.

WT1 as a complementary marker: WT1 is another transcription factor highly expressed in CIC-rearranged sarcomas. The paper shows that WT1 expression was elevated in the CIC-rearranged candidates but was not the highest overall among all test cohort samples. When ETV4 and WT1 are combined, they provide better discrimination than either alone, but the RF's use of 100-gene profiles outperformed any two-gene combination. The underlying biology supports this: CIC-rearranged sarcomas have a complex coordinated transcriptional program driven by the CIC::DUX4 fusion's activity as a transcriptional activator, and no single downstream target fully captures that program.

Clinical implications: In a diagnostic laboratory, ETV4 immunohistochemistry can serve as a rapid screening tool, but a positive result should prompt reflex molecular testing rather than immediate classification. The RF expression classifier, when run on existing RNA sequencing data, provides a more reliable second-level confirmation. The paper advocates for a tiered diagnostic algorithm: first fusion testing by targeted RNA-seq, then expression-based classification for fusion-negative cases with SBRCS morphology.

TL;DR: ETV4 expression above the 97th percentile flagged 15 true candidate CIC-rearranged cases but also captured 22 false positives in the 1,335-case cohort. The random forest's 100-gene model correctly down-ranked these false positives, demonstrating that coordinated multigene expression profiling substantially outperforms single-marker approaches.
Pages 10-12
Limitations, Caveats, and the Road to Clinical Deployment

The authors acknowledge several concrete limitations that contextualize how far the classifier is from routine clinical deployment. The most fundamental issue is sample size. The training cohort contained only 23 samples used to train the RF (13 Ewing sarcomas and 10 confirmed CIC-rearranged sarcomas). While the OOB cross-validation framework partially compensates for this, a larger and more diverse training cohort would reduce the risk of overfitting to institution-specific or cohort-specific expression patterns.

The CIC::DUX4 detection failure and its broader implication: The paper highlights a troubling finding: despite using a specialized Archer FusionPlex Sarcoma Kit on a subset of samples, CIC::DUX4 fusions remained undetectable in the 5 CIC-rearranged-like cases. The authors note this is likely due to the repetitive nature of the DUX4 gene locus on chromosome 4q35, which contains multiple near-identical copies that interfere with alignment and fusion calling algorithms. This failure underscores a limitation not just of this study but of any NGS-based RNA fusion assay for CIC-rearranged sarcomas, and it directly motivates the expression-based approach.

Single-center validation: All 1,404 samples (training plus test) came from a single institution in Erlangen, Germany. The classifier was trained and tested on the same platform (Illumina TruSight RNA Fusion Panel) under the same laboratory conditions. Whether the expression-based classifier generalizes to other institutions using different RNA sequencing panels, different FFPE processing protocols, or different bioinformatics pipelines is not established. Cross-center validation is a critical next step before the tool can be promoted for broad clinical use.

Future directions proposed: The authors identify three priority areas for future work. First, expanding the training cohort through multi-institutional collaboration to include more CIC-rearranged, BCOR-rearranged, and EWSR1::non-ETS fusion cases. Second, prospective functional or clinical follow-up of the 15 candidate cases to confirm their CIC-rearranged classification through additional orthogonal methods such as whole transcriptome sequencing or ATAC-seq. Third, extending the classifier framework beyond the CIC-vs-Ewing binary to a full multi-class SBRCS classifier covering all four Ewing-like subtypes simultaneously, which would require additional annotated training data but would have substantially greater clinical utility as a triage tool.

TL;DR: Key limitations include a small training cohort (23 samples), single-center validation using one sequencing platform, and the unresolved question of whether CIC-rearranged-like cases are truly CIC::DUX4-positive with a detection failure vs. a genuinely distinct entity. Priorities for future work are multi-institutional cohort expansion, orthogonal validation of the 15 candidate cases, and development of a full multi-class SBRCS classifier.