Full text 2026

A machine learning-derived genomic dataset from bacteria frequently reported as probiotics

Neres Rodrigues DL, Sodrzeieski PA, Auger S, et al.

Full text

Loading PDF… Expand reader Download

Abstract

Probiotics are live microorganisms that have been widely investigated for their association with beneficial host outcomes, particularly in the context of gut-associated microbial communities. Despite extensive literature, the probiotic effects are recognized as strain-specific and highly context-dependent, which limits the identification of universal genetic determinants of probiosis. In this study, we present a machine learning-derived genomic dataset generated from comparative analyses of bacterial genomes belonging to taxa frequently reported as probiotics and reference gut-associated bacteria. Using pangenomic analysis combined with supervised machine learning approaches, including Random Forest, Support Vector Machine, and Logistic Regression, we extracted discriminative genomic features from large-scale genome data. The resulting dataset comprises 1,072 non-redundant protein-coding sequences, accompanied by gene presence-absence matrices and functional annotations. These features should not be interpreted as causal determinants of probiotic functionality, but rather as genomic patterns associated with bacterial taxa commonly used as probiotics, which may also reflect taxonomic and ecological signatures. All data and scripts used in this study are publicly available through an open-access repository, providing a reusable resource for exploratory analyses, comparative genomics, and methodological benchmarking in probiogenomics and microbial genomics. The final data, hereby called ProbioSML, is currently available on https://doi.org/10.5281/zenodo.14181443.

Keywords

Bioinformatics data mining Gut Microbiota Probiogenomics Data Science