PVAED: prior-guided variational autoencoders with diffusion denoising for interpretable single-cell representation learning
Abstract
Single-cell RNA sequencing data rely heavily on dimensionality reduction methods to uncover underlying structures and patterns. Nonlinear dimensionality-reduction methods such as uniform manifold approximation and t-distributed stochastic neighbor embedding have been routinely used for this task. However, gene expression data alone often fails to capture and identify changes in cellular pathways, protein complexes, and TF-targets, which are more enlightening at the regulation level. To address this limitation, we present PVAED, a dimensionality reduction framework that integrates biological prior knowledge into a variational autoencoder (VAE) and further refines its latent embeddings with a diffusion-based denoising module. In addition, we incorporate a neighborhood-preserving loss term to ensure local similarity among cells in the reduced space. Across multiple benchmarks, PVAED achieves an average 43% improvement in low-dimensional cell representation compared to standard VAEs (evaluated across seven clustering metrics) and delivers a 34% (44%) gain in projection quality relative to classical or VAE-based approaches in terms of global (local) structure preservation. Beyond performance, PVAED offers improved interpretability through its prior-based encoder design, enabling prioritization and modeling of interpretable biological variables as hypergraphs. This framework facilitates the identification of critical regulatory factors associated with disease pathology and could also be utilized to predict potentially related genes for known pathways. Finally, we demonstrate the utility of PVAED by revealing pre-committed neuronal subpopulations in the differentiation of migrating neurons within the mammalian cerebral cortex.