Most computational biomarker discovery approaches identify genes based on their individual statistical association with a disease class, then present a ranked list as the result. This study argues that such an approach is fundamentally incomplete because it ignores the biological interactions between candidate markers. A gene with high discriminatory power in isolation may be functionally inactive in the tumor microenvironment if its upstream regulators or downstream effectors are also dysregulated.
The paper builds on prior work that identified a 96-gene panel capable of distinguishing small round blue cell tumors (SRBCTs) in children, a group of histologically similar pediatric cancers that includes rhabdomyosarcoma (RMS), Ewing's sarcoma (EWS), Burkitt lymphoma (BL), and neuroblastoma (NB). These tumors frequently masquerade as one another clinically, and misdiagnosis can direct patients toward ineffective treatment regimens.
The proposed framework, called Artificial Neural Network Inference (ANNI), uses backpropagation neural networks to model how the expression of each gene in the biomarker pool can be predicted from the expressions of all other genes. The resulting prediction weights and signal directions encode the strength and nature (stimulatory or inhibitory) of pairwise gene interactions, generating an interactome network map for the sarcoma biomarker set.
The dataset used was the publicly available SRBCTs cDNA microarray dataset originally published by Khan et al., which contains expression measurements for 2,308 genes across 88 samples. These samples span four pediatric tumor types: 25 RMS, 29 EWS, 11 Burkitt lymphoma (BL), 18 neuroblastoma (NB), and 5 unknown samples.
In this study, the analysis focused exclusively on the two sarcoma types, RMS and EWS, as these are the most biologically interrelated among the four groups. RMS originates from skeletal muscle progenitors and is most common in head/neck and genitourinary regions in children. EWS typically arises in bone and shares cytogenetic features with primitive neuroectodermal tumors (PNET), characterized by the recurrent EWS/FLI fusion on chromosome t(11;22)(q24;q12).
The 96-gene biomarker panel used as input to ANNI was previously identified using a hybrid genetic algorithm-neural network (GANN) model, which iteratively evaluated chromosomes representing gene subsets over 3,000 fitness evaluations repeated 5,000 times. Among the 96 GANN-selected genes, 44 were complementary to the top 96 genes originally reported by Khan et al., expanding the candidate set.
ANNI employs a 3-layer multilayer perceptron (MLP) with backpropagation learning and sigmoid activation function. The network architecture is N-2-1, where N equals the total number of genes minus 1. The core algorithm iterates through each gene in the panel: one gene at a time is designated as the output node, and all remaining genes serve as input nodes. The ANN learns to predict the output gene's expression from all others, and the resulting learned connection weights represent the interaction strengths between the output gene and each input gene.
Network training parameters included a learning rate of 0.1, momentum of 0.5, a maximum of 300 epochs, and a mean squared error threshold of 0.01 with a window of 100 epochs. Monte Carlo cross-validation (MCCV) was applied to prevent overfitting: samples were randomly partitioned 50 times into training/test/validation sets in a 60:20:20 ratio. The entire process was repeated 10 times per gene, and interaction scores were averaged across all repetitions.
Interaction directions were inferred from the sign of the Pearson correlation coefficient between the predicted interaction scores and the actual gene expression values, using a cutoff of r = 0.7. Interactions below this threshold were excluded from the final network map. Network visualization was performed using Cytoscape 2.8. On a simulated validation dataset of 100 samples and 25,000 features with 32 pre-defined correlating features, ANNI with 2 hidden nodes achieved a Pearson r of 0.805 between predicted and actual correlations, 89.16% correct sign assignment, and 93.16% true positive rate.
Analysis of all 96 genes produced a matrix of 9,120 potential interactions (96 x 95). To create an interpretable network, only the strongest single association for each of the 96 genes was retained, yielding 96 interaction edges for visualization. Within this simplified map, TNNT1 and FNDC5 emerged as the primary interaction hubs for RMS.
TNNT1 (troponin T1, primarily expressed in heart muscle) was strongly inhibited by RXRG, MYL1, RND3, FHL3, and FGFR4. FGFR4, a tyrosine kinase receptor known to be undetectable in normal tissue but activated in tumors, was found to be overexpressed in the RMS network and showed both direct inhibitory and stimulatory interactions with TNNT1. High FGFR4 expression has been previously associated with advanced-stage RMS and poor survival. The low activity of tumor suppressors FHL3, RXRG, and MYL1 may promote TNNT1 mutation, and the high expression of mutated TNNT1 (a diagnostic marker for nemaline myopathy) suggests that RMS tumors in this dataset are likely congenital.
FNDC5 (fibronectin domain containing protein 5, the precursor to irisin hormone) was inhibited by TNNT2, BIN1, SEPT4, MYL4, and HSPB2 in the RMS network. Overexpression of TNNT2, BIN1, and MYL4, all proteins involved in cardiac development and function, suggests that RMS patients in this dataset have lower risk of developing cardiomyopathy, because high expression of these cardiac proteins may suppress the FNDC5 pathway that promotes cardiomyopathic effects.
FCGRT (Fc fragment of IgG receptor and transporter), which acts as a promoter marker for the EWS-FLI fusion characteristic of Ewing tumors, was the primary interaction hub for EWS. Highly expressed FCGRT was suppressed by TLE2, CITED2, CAV1, PTTG1IP, and KDSR in the network. Notably, overexpressed CITED2, a cardiac transcription factor that inhibits hypoxia-induced gene activation, is known to induce resistance to cisplatin, a commonly used chemotherapeutic agent in sarcomas, suggesting a mechanism of drug resistance in EWS.
Overexpressed CAV1 (caveolin-1) acts as an additional promoter marker for EWS-FLI fusion and promotes metastasis in Ewing tumors. Low expression of TLE2 and PTTG1IP indicates promotion of tumor cell differentiation via disruption of the Runx2 osteoblast master regulator, which is normally blocked by EWS-FLI fusion, and histone modification via NKX2-2 (a known immunohistochemical marker for Ewing sarcoma).
The second EWS hub, OLFM1 (olfactomedin 1, highly expressed in brain), was suppressed by GYG2, APLP1, CHD3, TNFAIP6, and ARSB. This cluster revealed a triangular interaction between Wnt signaling, Fas/Rho signaling, and intracellular oxygen (ARSB) pathways. Overexpression of OLFM1 suppresses extracellular Wnt inhibitors, and deficiency of ARSB due to hypoxia leads to GAG accumulation in lysosomes. Overexpressed CHD3 (histone deacetylase complex component) further promotes tumor proliferation and differentiation.
ANNI produces a computational model of potential gene interactions, but these remain in silico predictions that require experimental validation in cell lines, animal models, or primary patient samples. The correlation-based interaction scores reflect statistical relationships in the gene expression data, which may not directly reflect actual protein-protein interactions, transcriptional regulation, or causal biological mechanisms.
The simplified network map (retaining only the strongest single association per gene out of 9,120 possible interactions) necessarily discards many potentially important secondary and tertiary interactions. Choosing a Pearson r cutoff of 0.7 to define significant interactions is an arbitrary threshold that may exclude biologically relevant lower-strength interactions.
The analysis is based on the SRBCTs dataset, which is relatively small and was collected under conditions that may not be representative of broader pediatric sarcoma populations. The algorithm was not applied to BL or NB in this publication, and validation of the identified interaction hubs (TNNT1, FNDC5, FCGRT, OLFM1) in independent sarcoma datasets is still required.
The ANNI framework provides a set of high-confidence in silico hypotheses for biological interactions that can drive experimental work. The cisplatin resistance mechanism mediated by CITED2 in EWS, for example, could be tested by knockdown experiments in EWS cell lines treated with cisplatin, quantifying changes in apoptosis and drug sensitivity. Similarly, FGFR4 inhibition studies in RMS models could validate its role as a hub node and potential therapeutic target.
Expanding ANNI to the full SRBCTs dataset including BL and NB, and applying it to other STS datasets, would allow cross-cancer comparison of gene interaction architectures. This could reveal whether certain hub genes are sarcoma-type-specific or shared across multiple pediatric tumor types, informing multi-cancer therapeutic strategies.
The ANNI algorithm source code was made available to collaborators at the time of publication. Integration with modern knowledge graphs, STRING protein interaction databases, and single-cell RNA-sequencing data could substantially extend the resolution and accuracy of the interaction maps, potentially enabling more precise identification of driver versus passenger interactions in these tumors.