Full text 2026

PZLAST-MAG: full length protein sequence similarity search server of large-scale MAG proteins

Higashi K, Ishikawa H, Kurokawa K, et al.

Full text

Loading PDF… Expand reader Download

Abstract

<h4>Motivation</h4>Metagenome-assembled genomes (MAGs) provide access to novel protein sequences from uncultured microbes, offering invaluable resources for studying protein diversity, structure prediction, and evolutionary analysis. However, despite the explosive growth of MAG-derived protein data, tools enabling fast and accurate similarity searches against large-scale MAG protein datasets remain limited.<h4>Results</h4>We present PZLAST-MAG, a web server for ultra-fast sequence similarity searches against 0.4 billion MAG-derived protein sequences (0.1 trillion amino acids) from over 210 000 MAGs indexed in Microbiome Datahub. Implemented on PEZY-SC3 MIMD many-core processors, PZLAST-MAG achieves high accuracy and speed, with performance comparable to widely used tools such as DIAMOND and MMseqs2 based on our benchmark analyses. In addition to tabular alignments, PZLAST-MAG provides interactive visualizations of phylogenetic and environmental distributions and co-occurrence patterns of homologous proteins across MAGs. This combination enables rapid homolog mining of functionally important genes across diverse microbial lineages while simultaneously revealing their taxonomic and ecological contexts. Two use case analyses indicate its utility for homolog mining of metabolic enzyme genes and plasmid-derived genes.<h4>Availability and implementation</h4>PZLAST-MAG is provided as a web-based service and is freely available at https://pzlast.nig.ac.jp/pzlast/mag without requiring registration.