A probabilistic approach for predicting indole-3-acetic acid synthesis in bacteria using genomic data
Abstract
BACKGROUND: Indole-3-acetic acid (IAA) is a crucial plant hormone for growth and development. Among bacteria capable of IAA synthesis, tryptophan-dependent pathways are the primary pathway for IAA production. Identifying IAA-producing bacteria is essential for understanding plant–microbe interactions and developing sustainable agricultural practices. With the advancements in sequencing technologies which enable the acquisition of genomic data in a short amount of time at a reduced cost. This study developed a probabilistic framework to predict the likelihood of IAA synthesis in bacteria using their genomes, which offers a cost-effective screening approach compared to traditional wet-lab methods. RESULTS: The prediction framework incorporates four tryptophan-dependent pathways and consists of two models. The Binary Model provides a binary prediction of IAA synthesis, and the Logistic Model ranks bacteria based on their synthesis potential. The models were validated against a comprehensive dataset of more than 15,000 bacterial genomes, and the results suggest that only 1.7% of the bacteria have the potential to synthesize IAA. The most promising phyla are Actinobacteria, Proteobacteria, and Firmicutes. The Logistic Model shows that the top-ranked genera include Streptomyces, Paraburkholderia, Nonomuraea, Rhodococcus, and Nocardia. The model also suggests that the potential for IAA synthesis is likely to be strain-specific, emphasizing the importance of analyzing individual strains within a genus. CONCLUSION: The proposed probabilistic prediction framework is a cost-effective in-silico approach for identifying bacteria with a high potential for IAA synthesis. This serves as a pre-screening tool for resource-intensive downstream wet-lab validation, significantly reducing the time and effort required to identify potential plant growth-promoting bacteria. Furthermore, the model can be customized for other pathways of interest, and making it a versatile framework to explore various bacterial metabolic processes. The source code of the model is freely available at https://github.com/ComputationalAgronomy/bcpip