Classification and Diagnostic Prediction of Cancers Using Gene Expression Profiling and Artificial Neural Networks

Nat Med 2001 AI 8 Explanations View Original
Original Paper (PDF)

Unable to display PDF. Download it here or view on PMC.

Plain-English Explanations
Pages 1-2
Why SRBCT Diagnosis Remains Difficult Without Molecular Methods

The small round blue cell tumors (SRBCTs) of childhood, including neuroblastoma (NB), rhabdomyosarcoma (RMS), non-Hodgkin lymphoma (NHL, specifically Burkitt lymphoma), and the Ewing family of tumors (EWS), are named for their identical microscopic appearance under standard light microscopy. Despite being four biologically distinct cancers, they cannot be reliably distinguished by histology alone. Yet treatment protocols, drug responses, and prognosis vary dramatically across these four categories.

Current clinical practice uses several supplementary diagnostic techniques: immunohistochemistry for specific proteins (limited to one protein per test), cytogenetics for chromosomal translocations, fluorescence in situ hybridization (FISH), and RT-PCR to detect tumor-specific fusion genes such as EWS-FLI1 in Ewing sarcoma (t(11;22)(q24;q12)) and PAX3-FKHR in alveolar rhabdomyosarcoma (t(2;13)(q35;q14)). However, these individual tests can fail: technical difficulties, variant translocations, and the fact that some markers like MIC2 appear in multiple cancer types all limit specificity.

Gene expression profiling using cDNA microarrays measures thousands of genes simultaneously, providing a comprehensive molecular fingerprint of each tumor. However, no prior study had rigorously tested multiclass diagnostic classification using microarray data for more than two cancer categories simultaneously. This study is the first application of artificial neural networks (ANNs) trained on microarray data for multiclass cancer classification, and the first to systematically identify the specific genes driving each classification decision.

TL;DR: SRBCTs look identical under a microscope but require completely different treatments. Existing tests (immunohistochemistry, FISH, RT-PCR) are single-marker and can fail. This is the first ANN study to classify 4 cancer types from microarray data and identify key genes, published in Nature Medicine 2001.
Pages 2, 6, 7
PCA Dimensionality Reduction and Three-Fold Cross-Validation with 3,750 Models

Starting with 6,567 genes from cDNA microarrays, a quality filter requiring red intensity greater than 20 across all experiments reduced the set to 2,308 genes. Principal component analysis (PCA) was then applied to reduce dimensionality further: the 10 dominant PCA components, capturing 63% of the total variance across all samples, were used as inputs to the ANN models. The output layer had four nodes, one for each cancer category (EWS, RMS, NB, BL), producing values between 0 and 1.

A rigorous three-fold cross-validation was used to avoid overfitting. The 63 training samples were randomly split into three equally-sized groups; two groups calibrated the model and one validated it, rotating through all three possibilities. This was repeated with 1,250 different random splits, producing a total of 3,750 calibrated ANN models. Final classification of any sample used the committee vote: the average output across all 3,750 models. Each sample's committee vote was compared to the ideal output for each class, and only samples within the 95th percentile of the training distribution were assigned a confident diagnosis.

To identify genes driving classification, the sensitivity of each ANN model's output to small perturbations in each gene's expression level was computed as the absolute value of the partial derivative of the output with respect to each gene, averaged across all samples and all 3,750 models. A large sensitivity means that small changes in that gene's expression substantially change the classification outcome. Genes were ranked by this sensitivity score, and the minimum gene set achieving zero misclassification was determined by progressively adding higher-ranked genes.

TL;DR: 2,308 genes pass quality filter from 6,567. PCA reduces to 10 components capturing 63% of variance. 3,750 ANN models from 1,250 random 3-fold splits. Gene importance = sensitivity of output to expression perturbation. Minimum error-free gene set: 96 genes.
Pages 2-3
100% Classification Accuracy on Blinded Test Samples Using 96 Genes

Using all 3,750 ANN models, all 63 training samples were correctly assigned to their respective categories. When calibrated using the top 96 ranked genes (which captured 79% of variance in the data matrix), the models again achieved perfect classification of all 63 training samples. Multidimensional scaling (MDS) analysis of these 96 genes showed clear, well-separated clusters for each of the four cancer types.

On the 25 blinded test samples, the ANN committee correctly classified all 20 SRBCT test cases with the highest vote going to the correct category in every instance. The diagnostic confidence threshold (95th percentile of training distance distribution) was used to assess confidence: 3 SRBCT samples (EWS-T13, Test 10, Test 20) were correctly classified but fell outside the 95th percentile, meaning diagnosis was correct but flagged as uncertain. All 5 non-SRBCT samples (2 normal muscle, 1 undifferentiated sarcoma, 1 osteosarcoma, 1 prostate carcinoma) were correctly rejected from all four categories.

Sensitivity and specificity by cancer type for all 88 samples: EWS sensitivity 93%, RMS sensitivity 96%, NB sensitivity 100%, BL sensitivity 100%. Specificity was 100% for all four categories. Hierarchical clustering using the 96 genes independently confirmed the ANN results: all 20 SRBCT test samples clustered within their correct diagnostic categories, and the two BL cell line replicates (ST486 C2 and C4) as well as the two NB replicates (GICAN C2 and C7) each clustered adjacent to one another.

TL;DR: 96 genes achieve 0% misclassification on 63 training samples. On 25 blind test samples: 20/20 SRBCTs classified correctly, 5/5 non-SRBCTs rejected. Sensitivity: EWS 93%, RMS 96%, NB 100%, BL 100%. Specificity: 100% for all four classes.
Pages 2, 5, 6
Training Set Design: Tumor Biopsies and Cell Lines Combined

The 63 training samples included both tumor biopsy material (13 EWS and 10 RMS tumors) and cell lines (10 EWS, 10 RMS, 12 NB, 8 Burkitt lymphomas). This design was deliberate: tumor tissue reflects gene expression in the context of the in vivo environment (including stromal cells), while cell lines provide a pure malignant population without stromal contamination. The combination compensates for weaknesses in each type alone.

Two pairs of samples (ST486 for BL and GICAN for NB) were each independently run twice on the microarrays and treated as separate samples, providing an internal reproducibility check. About 20% of samples from each category were randomly withheld as test samples; 4 additional NB tumors and 5 non-SRBCT samples were added to the test set blinded to the analysis team, for a total of 25 test samples.

All original diagnoses were made at tertiary hospitals with reference diagnostic laboratories experienced in pediatric cancers. The EWS samples covered the expected range of EWS-FLI1 translocations; RMS samples included both alveolar RMS (with PAX3-FKHR translocation) and embryonal RMS; NB samples included both MYCN-amplified and single-copy cases. The diversity of molecular subtypes within each histological category strengthens the generalizability of the ANN models.

TL;DR: 63 training samples: tumor biopsies and cell lines combined for complementary coverage. 4 different institutions contributed samples spanning molecular subtypes. Test set: 25 samples including 5 non-SRBCTs, blinded to analysts. Internal reproducibility confirmed by 2 duplicate cell-line pairs.
Pages 3-5
Ninety-Six Discriminating Genes and Their Biological Meaning

The 96 ranked genes include 93 unique genes (IGF2 represented by 3 clones, MYC by 2). Among them: 16 genes are specifically highly expressed in EWS, 20 in RMS, 15 in NB, and 10 in BL. Twelve genes discriminate primarily by their lack of expression in BL. One gene (EST Clone ID 295985) discriminates EWS by being specifically absent from that cancer type. Of all 61 genes specifically expressed in a single cancer type, 41 had not previously been reported as associated with these diseases.

FGFR4 (fibroblast growth factor receptor 4) emerged as a top RMS-specific marker. It is expressed during myogenesis but not in adult muscle, and is of interest for tumor growth and prevention of terminal muscle differentiation. Immunohistochemistry on SRBCT tissue arrays confirmed moderate to strong cytoplasmic staining for FGFR4 in all 26 RMS cases tested (17 alveolar, 9 embryonal), with only weak staining in EWS and NHL. In NB, the neural-specific genes TUBB5, ANXA1, NOE1, and GSTM5 appeared prominently, lending support to the proposed neural origin of EWS.

A key finding is that ANN-based gene ranking can identify biologically relevant genes without requiring them to be exclusively associated with a single cancer type. The method ranks genes by their contribution to the overall multiclass discrimination, which allows genes expressed in two of four categories to still be selected if the pattern provides discriminating power. This is more reflective of real tumor biology, where lineage markers are often shared across related tumor types.

TL;DR: 96 genes: 16 EWS-specific, 20 RMS-specific, 15 NB-specific, 10 BL-specific. 41 of 61 type-specific genes were novel discoveries. FGFR4 confirmed by immunohistochemistry in all 26 RMS cases. Neural genes support EWS neural origin. ANN finds genes that don't need to be exclusively single-type markers.
Pages 3-4
Committee Voting, Confidence Thresholds, and Handling Uncertain Diagnoses

The committee voting approach across 3,750 models provides both a classification decision and a confidence measure. For each sample, the committee vote for each of the four categories is computed, and the category with the highest vote determines the classification. The distance from the committee vote to the ideal vote (e.g., EWS = 1, RMS = NB = BL = 0) measures classification confidence: a smaller distance means higher confidence. If this distance exceeds the 95th percentile of the training distribution for that category, the diagnosis is rejected.

This threshold-based rejection mechanism is clinically important: in practice, some tumors may be non-SRBCT, undifferentiated, or outside the four categories the model was trained on. All 5 non-SRBCT test samples fell outside all four 95th percentile boundaries, meaning the system correctly refused to assign any diagnosis. Three SRBCT samples were correctly classified by highest vote but had distances exceeding the 95th percentile, flagging them for additional clinical investigation rather than making a confident but potentially incorrect diagnosis.

The use of Support Vector Machines (SVMs) had been proposed by others for similar problems, but prior SVM studies had not been tested for the ability to identify key discriminating genes. The ANN sensitivity analysis provides a natural, principled way to extract gene importance, a feature that SVMs do not inherently provide without additional methods. Hierarchical clustering of the 96 ANN-identified genes as an independent method reproduced the same classification structure, validating that these genes capture real biological signal.

TL;DR: 3,750 models vote per sample. Distance from ideal vote provides confidence. Samples beyond 95th percentile are rejected rather than misclassified. 5/5 non-SRBCTs correctly rejected. 3 borderline SRBCTs flagged as uncertain. ANN provides gene importance scores that SVM methods cannot easily extract.
Pages 4-5
Clinical Applications: Subarrays, Prognosis, and Therapy Targets

Establishing that 96 genes can achieve zero misclassification opens a practical path to cost-effective clinical diagnostics. A targeted SRBCT subarray containing only these 96 genes (or potentially even fewer once validated in prospective studies) could replace full-genome arrays, dramatically reducing cost and complexity. The 93 unique discriminating genes represent candidates for such a custom panel.

Several genes in the 96-gene set, particularly those not previously reported in these cancers, represent potential therapeutic targets. For example, FGFR4's strong and consistent expression in RMS makes it an attractive target for receptor tyrosine kinase inhibitors in that cancer type. The 41 novel type-specific genes identified in this study provide a rich set of candidates for follow-up biological and therapeutic studies.

The authors explicitly envision extending the method beyond diagnostic classification to predict tumor stage, prognosis, and treatment response. By training ANNs on samples with known clinical outcomes rather than just histological category, the same methodology could identify gene signatures predictive of metastasis, relapse, or drug sensitivity. This extension would directly address clinical decision-making beyond initial diagnosis.

TL;DR: 96-gene SRBCT subarray could enable cost-effective clinical diagnostics. FGFR4 is a validated RMS therapeutic target. 41 novel type-specific genes are candidates for drug development. Method extension to prognostic and treatment-response prediction is the proposed next step.
Page 4
Training Size Constraints, Stromal Contamination, and Generalizability

A fundamental limitation is the small training dataset: 63 samples across four categories means an average of fewer than 16 samples per class. This constraint drove the decision to use linear ANN models (no hidden layers) rather than nonlinear ones, to limit the number of free parameters relative to the sample count. The authors acknowledge that nonlinear features could improve classification if more data were available, and expect future work with larger cohorts and more comprehensive arrays to improve sensitivity beyond the current levels.

Tumor biopsy samples contain a mixture of malignant cells and stromal cells (fibroblasts, immune cells, vasculature), so expression levels from tumor tissue reflect a composite signal. This heterogeneity is partially compensated by also training on pure cell lines, but it means that some genes identified as markers may reflect stromal differences between tumor types rather than intrinsic tumor biology. The relative contributions of tumor versus stromal expression to the 96 discriminating genes were not independently assessed.

The stringent quality filter applied before analysis removes genes with missing or low-quality measurements across all samples. This improves model robustness but may exclude some truly important genes that happen to show artifact in one sample. The authors note that expanding to larger, more comprehensive arrays would likely extend the 96-gene list and potentially reveal additional biology.

TL;DR: Only 63 training samples (average 16 per class) forced use of linear ANN models. Tumor biopsies contain stromal contamination. Quality filtering may exclude some true markers. Larger arrays and cohorts expected to improve sensitivity beyond current EWS 93%, RMS 96% levels.
Citation: Open Access, . Available at: PMC1282521.