USADAE: a deep learning approach to disentangle hidden covariates in RNA-seq data
Abstract
Integrative analysis of RNA-seq datasets faces critical challenges in disentangling biologically meaningful signals from implicit confounders, such as unmeasured technical variability and nonlinear interactions between biological and technical variables. Traditional methods like Surrogate Variable Analysis (SVA), Probabilistic Estimation of Expression Residuals (PEER), and Remove Unwanted Variation (RUVSeq) rely on linear assumptions, which generally fail under complex nonlinear confounding patterns. Although deep learning approaches show promise in single-cell RNA-seq, they primarily address known batch labels rather than disentangling hidden confounders. Here, we develop a novel autoencoder framework coupled with adversarial learning, UnSupervised Adversarial Deconfounding AutoEncoder (USADAE), specifically designed to separate confounders from biological signals. The model encodes RNA-seq data into distinct biological and confounder latent spaces through adversarial disentanglement, enabling downstream correction of differential expression and eQTL analysis. In comprehensive simulations, USADAE significantly outperforms existing methods in extracting covariates while preserving biological signals. Real-data applications further demonstrate its robustness across diverse scenarios, including cancer genomics and eQTL studies.