Unlocking the Next Decade of Proteomics with Standardized, Structured Metadata
Abstract
The proteomics community has fully embraced data sharing, yet data set metadata provision remains limited, especially at the level of the biological samples and experimental design. This hampers large-scale data reuse, as comprehensive and structured sample context and study design information are often essential for confident, automatic reuse, and (re)interpretation. Although standards such as Sample and Data Relationship Format for Proteomics (SDRF-Proteomics) and supporting tools are already available, their adoption remains limited. Many researchers lack incentives, and enforcement by journals and repositories remains challenging in practice. Still, metadata defines a data set's long-term value. We propose a coordinated plan to dramatically improve metadata annotation of publicly disseminated proteomics data. Funders can drive progress by investing in a sustainable, scalable metadata infrastructure. HUPO-PSI plays a central role in setting community standards and enabling validation. ProteomeXchange repositories are key to implementing and supporting metadata adoption. Data producers must treat metadata as a part of their scientific output. Instrument vendors can contribute by enabling the automatic capture of technical metadata. Software developers should embed SDRF-Proteomics metadata into analysis workflows. Finally, journals and reviewers are well positioned to shape expectations and enforce compliance. By aligning efforts across these stakeholders, we can build the road to large-scale, context-aware reuse and unlock the full value of public proteomics data sets.