Full text 2026

PyTEA-O: a Python implementation of Two-Entropies Analysis for protein sequence variation analysis

Kuin RCM, Julian AT, Chander J, et al.

Full text

Loading PDF… Expand reader Download

Abstract

<h4>Motivation</h4>Protein sequence variation analysis is a topic of broad interest in drug discovery and protein engineering to support modulation of protein function for diverse biotechnological and therapeutic applications. To assist in the analysis of multiple sequence alignments (MSAs) and identify residues that account for protein function specificity, computational tools have been developed. Yet, existing programs often omit consideration of amino acid properties, flexibility beyond fixed webserver interfaces, accessible source code, or compatibility with small MSAs.<h4>Results</h4>To address these limitations, we present PyTEA-O, a Python implementation of Two-Entropies Analysis that has been developed to be easy to use for the analysis of protein sequence variation. To help users analyze the MSA and screen for residues of interest, we generate modifiable and intuitive visualizations. These visualizations, together with a scoring approach for identifying alignment positions with (dis-)similar physicochemical properties, presents a powerful tool for sequence variability analysis. To demonstrate its capabilities, we present a case study based on the deubiquitinase OTUD7B (Cezanne) where we identify a crucial position that modulates its affinity for its substrate.<h4>Availability and implementation</h4>PyTEA-O is available at https://github.com/CDDLeiden/PyTEA-O/ and archived via Zenodo (https://doi.org/10.5281/zenodo.15914598).