CanID: A Robust and Accurate RNA-seq Expression-based Diagnostic Classification Scheme for Pediatric Malignancies
Abstract
Cancer subtype classification is critical for precision therapy, and there is a growing trend to augment histopathology testing with omics-based machine learning classifiers. However, analytical challenges remain in pediatric cancer regarding the scope and precision of current classifiers, as well as the evolving subtype standardization. To address these challenges, we constructed Cancer Identification (CanID), a stacked ensemble machine learning classification scheme, using transcriptomic features derived from gene-level RNA sequencing count data as the sole input. CanID was developed primarily from 3203 pediatric cancer samples across 13 solid tumor subtypes and 38 hematologic malignancy subtypes, with subtype labels curated without the use of RNA-seq data. The accuracies of independent testing in three independent or external datasets for solid tumors and hematologic malignancies were 99% and 92%-93%, respectively. Notably, CanID was able to classify subtypes challenging for clinical histology evaluation and was robust to both biological and technical challenges, including differences in data collection protocols, class imbalance, potential mislabeled training samples, and classes unobserved during training. The high accuracy, robustness, and biological interpretability of this transcriptome-based classification scheme represent a valuable approach to advance tumor diagnosis and clinically meaningful stratification of tumor types. CanID can be accessed on GitHub at https://github.com/chenlab-sj/CanID.