Soft tissue sarcoma (STS) is a rare malignancy arising from mesenchymal tissues, which include fat, muscle, nerves, fibrous tissue, and blood vessels. It is not one disease but a family of over 50 distinct subtypes, each with its own molecular signature, histological appearance, and clinical behavior. This heterogeneity makes STS one of the most diagnostically challenging cancer types, and that challenge extends directly into how patients are tracked and studied using administrative health claims databases.
Real-world data (RWD) sources such as insurance claims databases have become indispensable for oncology research. They capture large patient populations and long follow-up periods that prospective trials cannot match. However, because these databases rely on billing codes rather than confirmed clinical diagnoses, the accuracy of patient selection depends heavily on how well those codes reflect actual disease. For STS specifically, the International Classification of Diseases Clinical Modification (ICD-CM) system provides codes that can be misapplied or shared with other conditions, leading to misclassification.
The selection bias problem: A 2013 review of 64 oncology studies using administrative claims found that only 7% used a validated algorithm to select patients, 36% relied on a single ICD-CM diagnosis code, and only 5% even discussed how their selection criteria could influence their findings. This is a fundamental problem for any research conclusions drawn from such studies, because the study population may not accurately represent true STS patients.
This paper from Princic, McMorrow, Chan, and Hess (funded by Eli Lilly) set out to systematically test whether published identification algorithms, plus newly developed ones, could reliably select STS patients from a large linked claims-EMR database. The study was the first of its kind specifically for STS, and it arrived at a sobering conclusion: none of the 14 algorithms tested achieved both adequate sensitivity and specificity simultaneously.
The study used the IBM MarketScan Explorys Linked Claims-Electronic Medical Record (EMR) Dataset, known as the CED. This is a unique resource that links two independent data streams: the MarketScan Commercial Claims and Encounters Database and the Explorys EMR Database. MarketScan captures the inpatient, outpatient, and prescription drug experience of approximately 198.9 million employees and their dependents across a variety of employer-sponsored insurance plans. Explorys provides structured clinical data from integrated health networks.
Coverage and linkage: The combined CED contains approximately 4.5 million linked patients with both claims and EMR records covering the period from January 1, 2000 through July 31, 2018. The linkage is probabilistic, using patient demographics and other non-identifiable characteristics to connect records across the two systems. Patients in the CED therefore have a richer longitudinal history than either database could provide alone, including both the clinical detail from EMR records (diagnoses from treating physicians) and the complete claims trail of visits, procedures, and prescriptions.
Why linkage matters: The key methodological insight here is that the EMR record serves as the gold standard for true STS diagnosis, while the claims data is what investigators would normally work with in a pure claims study. By having both, the researchers could define true STS cases using clinical terminology coded in Systematized Nomenclature of Medicine (SNOMED) terms from physician records, and then test how well various claims-based algorithms captured those same patients.
All data were de-identified and HIPAA-compliant, exempting the study from IRB review. The study period spanned nearly 18 years, providing a substantial observation window for both cancer diagnosis coding patterns and treatment trajectories.
Two patient populations were identified from the CED: STS cases and non-STS cancer controls. STS cases were defined by the presence of a SNOMED-coded STS diagnosis on a clinical record in the Explorys EMR. Non-STS controls were patients who had any cancer diagnosis in the EMR but no evidence of STS on any clinical record. Both groups were required to be at least 18 years old and enrolled in administrative claims during a qualifying observation period.
Cohort sizes and demographics: After eligibility screening, there were 784 STS cases and 249,062 non-STS cancer controls available for analysis. STS cases were younger on average (mean age 59.6 years, SD 14.8) compared to controls (mean age 64.2 years, SD 12.7). The gender distribution was roughly equal in both groups, with slightly more females (54.2% of STS cases, 54.0% of controls). Among STS cases, 19.2% were aged 45-54, 30.7% were aged 55-64, and 16.4% were aged 65-74.
Split-sample validation: The dataset was divided into development and validation samples of roughly equal size. The development sample included 10,906 STS case panels and approximately 1.8 million control panels; the validation sample had 10,840 STS case panels and another 1.8 million control panels. A panel-based approach was used: for each patient, a panel was constructed around each eligible cancer diagnosis date (the index date), extending to disenrollment in claims or end of study. This allowed algorithms to be tested at multiple time points per patient.
Variable extraction: All algorithm variables were drawn from claims records and included imaging codes (CT, MRI, radiograph, PET), surgical procedure codes (excision and resection), symptom codes (pain in limb, neoplasm of uncertain behavior in skin, localized superficial swelling), cancer site codes, and treatment codes. No clinical free-text or pathology data were used, reflecting the constraints of a claims-only study.
The 14 algorithms represented a spectrum of complexity, from requiring only two ICD-CM STS diagnosis codes to demanding specific drug regimens combined with multiple clinical criteria. They were derived partly from previously published literature and partly from novel combinations constructed by the research team. All algorithms were applied to the diagnostic period, defined as the window surrounding each index cancer diagnosis date.
Algorithm 1 (the simplest baseline): Required at least two medical claims with an ICD-CM STS diagnosis code at least 30 days apart in any position. This mirrors what many studies use by default, a repeated code over time to reduce the chance that a single erroneous billing entry drives patient classification.
Algorithm 2 (adding treatment confirmation): Built on Algorithm 1 but additionally required at least one claim for an NCCN-recommended systemic therapy for STS following the first STS diagnosis. The NCCN treatment list included single agents such as doxorubicin, ifosfamide, epirubicin, gemcitabine, dacarbazine, liposomal doxorubicin, temozolomide, vinorelbine, eribulin, trabectedin, pazopanib, regorafenib, and larotrectinib, as well as combination regimens such as AIM (doxorubicin, ifosfamide, mesna), MAID (mesna, doxorubicin, ifosfamide, dacarbazine), and gemcitabine with docetaxel or vinorelbine.
Algorithms 3-6 (alternative criteria sets): These varied combinations of cancer site codes, treatment types, symptoms, procedures, and diagnosis timing. Some required any solid tumor code plus NCCN treatment without a specific STS code. Others incorporated imaging procedures or surgical codes. Algorithms 5a through 5e tested variants of a two-diagnosis requirement, varying the time gap between claims and whether they needed to appear in the primary versus any diagnosis position. Algorithm 6 series combined non-STS cancer codes plus specific symptoms or procedures.
The core finding of this study is stark: none of the 14 tested algorithms simultaneously achieved sensitivity and specificity in a range considered acceptable by prior oncology claims research standards. The benchmark was set using published literature, which reported both sensitivity and specificity ranging from 73% to 95% for other cancer types. For STS, every algorithm fell outside this zone on at least one dimension.
Algorithm 1 performance: The simplest algorithm, requiring two ICD-CM STS codes at least 30 days apart, achieved a sensitivity of 59.0% and a specificity of 79.7%. The positive predictive value (PPV) was 43.2% and the negative predictive value (NPV) was 88.1%. In practice, a 59% sensitivity means roughly 4 in 10 true STS patients would be missed. A 43% PPV means more than half of patients flagged as STS would not actually have the disease.
Treatment-requiring algorithms: Algorithms that required NCCN-recommended pharmacologic treatment as a confirmatory criterion (Algorithms 2, 3, 4, 5b, and 6b) achieved high specificity of 91% to 99%, but their sensitivity collapsed to below 20%. The reason is straightforward: only approximately 24.5% of confirmed STS cases in the claims database had any record of NCCN-recommended systemic therapy following their index diagnosis. Requiring a treatment that most patients never receive in claims data guarantees high specificity but disqualifies the majority of true cases.
Intermediate algorithms: Variations that added imaging codes, symptom codes, or alternate cancer site codes alongside the STS diagnosis requirement did not meaningfully improve the sensitivity-specificity trade-off. No combination of these variables broke through the ceiling imposed by the fundamental unreliability of STS coding in claims data.
The poor algorithm performance is not arbitrary. It reflects structural features of STS as a disease that make it fundamentally difficult to capture accurately through billing codes. Understanding these root causes is important both for interpreting this study and for appreciating why machine learning approaches might be needed in the future.
Anatomical diversity: STS arises in soft tissues throughout the entire body. Approximately 43% of cases occur in the limbs, 19% in the gastrointestinal tract, 15% in the retroperitoneum, 10% in the trunk, and 9% in the head and neck. This spread means there is no single anatomical signature in claims data. Imaging studies, surgical procedures, and even symptoms will vary widely by site. A retroperitoneal STS may generate procedure codes that look like colorectal or renal cancer workup. A limb STS may generate codes indistinguishable from benign soft tissue masses or lipomas.
Non-specific symptoms and workup: STS typically presents with a painless mass or swelling that may be present for months before diagnosis. The symptoms (pain in limb, localized swelling, mass of uncertain behavior) and the standard diagnostic workup (MRI, CT, PET, core needle biopsy, excision) are shared with many benign and malignant conditions. There is no STS-specific biomarker that appears in claims data. This means the pre-diagnostic claims trail for an STS patient looks similar to the trail for someone with a benign lipoma or an inflammatory mass.
Undertreatment in claims: Only 24.5% of STS cases had evidence of NCCN-recommended systemic therapy in their claims record. This reflects the realities of sarcoma treatment: early-stage disease is often managed with surgery alone, and many patients receive institutional or off-label regimens not captured by standard NCCN lists. Relying on chemotherapy as a confirmatory criterion therefore excludes the majority of surgically managed patients, creating severe sensitivity loss.
The authors identify several important limitations that affect how broadly these results can be applied. The most fundamental is the nature of the gold standard itself. The study defined true STS cases using SNOMED-coded diagnoses in the Explorys EMR, but EMR records are only as complete as the clinical encounters they capture. Patients receiving care outside the networked clinics would not have STS documented in the EMR even if they have the disease. This means some patients classified as non-STS controls may actually have STS managed elsewhere, introducing misclassification in both directions.
Claims-EMR linkage limitations: The probabilistic linkage between MarketScan claims and Explorys EMR data means that some patients in the final linked dataset may not be correctly matched. Mislinked records would place a patient's claims history with the wrong EMR record, generating spurious case or control classifications. The researchers used standard quality checks but could not fully eliminate linkage error.
Population representativeness: The MarketScan database primarily represents commercially insured individuals, skewing toward working-age adults with employer-sponsored coverage. Older patients on Medicare and lower-income patients on Medicaid, who make up a significant share of the real STS population, are underrepresented. The mean age of STS cases in this study (59.6 years) suggests some skew toward younger patients relative to the general population distribution.
ICD coding evolution: The study period spanned the transition from ICD-9-CM to ICD-10-CM in October 2015. ICD-10-CM introduced more granular STS codes but also changed coding patterns. Algorithms calibrated on ICD-9 coding patterns may perform differently in the ICD-10 era, and the transition itself may have introduced inconsistencies in the data.
Temporal scope of treatment data: The observation period for algorithm testing was defined from the index cancer diagnosis date forward. Treatments administered before the index date (pre-diagnostic chemotherapy, neoadjuvant therapy) were not counted in the confirmatory criteria, which may have caused some cases to fail treatment-based algorithms even though they did receive appropriate STS-directed therapy.
The study's authors close by acknowledging that conventional algorithm approaches have likely hit a ceiling for STS identification in claims data. The combinations of diagnosis codes, treatment codes, imaging procedures, and symptom codes that can be assembled from ICD-CM billing records do not carry enough signal to reliably distinguish STS from the many other conditions that produce similar claims patterns. The authors explicitly recommend machine learning as the next investigative step.
The machine learning proposition: Rather than applying fixed, rule-based criteria in a specified time window, machine learning approaches can examine the timing, sequence, and co-occurrence patterns of hundreds of claims variables simultaneously. A gradient boosting or deep learning model trained on a linked claims-EMR dataset like the CED could, in principle, learn that a particular combination of MRI codes for a lower-extremity mass, followed by a specific surgical procedure code, followed by a specific pathology billing code, is highly predictive of STS even without an explicit ICD-STS diagnosis code appearing. The model can weight these patterns using labeled training data, generalizing rules that no human investigator would think to specify manually.
Implications for real-world evidence studies: Until better identification methods are available, researchers using administrative claims to study STS populations should be explicit about the limitations of their selection algorithms, report sensitivity and specificity estimates where possible, and acknowledge that selection bias is likely. The authors' benchmark finding that only 7% of prior oncology claims studies used validated algorithms suggests this is rarely done in practice. This paper provides a quantitative basis for the concern that STS populations in claims studies may be severely misclassified.
Broader applicability: The methodological challenge documented here is not unique to STS, but STS represents an extreme case because of its rarity, heterogeneity, and symptomatic overlap with benign conditions. The linked claims-EMR study design used here, with SNOMED-coded EMR diagnoses as the gold standard, offers a replicable framework for validating identification algorithms in other rare cancers where the same problems apply. The split-sample development and validation design also provides a template for how future algorithm studies should be structured to avoid overfitting.