Full text 2025

Proteomics Data Imputation With a Deep Model That Learns From Many Datasets

Harris L, Noble WS.

Full text

Loading PDF… Expand reader Download

Abstract

Missing values are a major challenge in the analysis of mass spectrometry proteomics data. Missing values hinder reproducibility, decrease statistical power for identifying differentially abundant proteins, and make it challenging to analyze low-abundance proteins. We present Lupine, a deep learning-based method for imputing, or estimating, missing values in quantitative proteomics data. Lupine is, to our knowledge, the first imputation method that is designed to learn jointly from many datasets, and we provide evidence that this approach leads to more accurate predictions. We validated Lupine by applying it to tandem mass tag data from >1000 cancer patient samples spanning 10 cancer types from the Clinical Proteomics Tumor Atlas Consortium. Lupine outperforms the state of the art for proteomics imputation, uniquely identifies differentially abundant proteins and Gene Ontology terms, and learns a meaningful representation of proteins and patient samples. Lupine is implemented as an open-source Python package.

Keywords

Mass spectrometry Proteomics Machine Learning Imputation Deep Learning