Full text 2026

ANOMALY: a Snakemake pipeline for identifying NuMTs from long-read sequencing data

Mahar NS, Singh R, Gupta I, et al.

Full text

Loading PDF… Expand reader Download

Abstract

Nuclear mitochondrial DNA segments (NuMTs) can contribute to cancer development and disease progression by disrupting protein-coding genes. Furthermore, their presence confounds mitochondrial variant detection, underscoring the critical need for robust NuMT detection. Current methods to call NuMTs rely on short-read sequencing data but struggle to resolve complex NuMTs. These limitations can be overcome by employing long-read sequencing data. However, no such workflow exists to capture NuMTs from long-read sequencing data. Here, we introduce ANOMALY, a novel, easy-to-use workflow for detecting NuMTs from long-read sequencing data. The pipeline takes raw sequencing or aligned data and calls and visualises sample NuMTs. On 50 simulated datasets, the pipeline demonstrated high accuracy, with a precision of 1.000, a recall of 0.989, and an F1-score of 0.994. The pipeline underscores the limitations of short-read data in resolving and capturing complex NuMTs while demonstrating that long-read data enables their accurate identification. The Snakemake pipeline employs Python, Bash and R and is published under an open-source GNU GPL v3 license. Detailed information on setting up and running the pipeline, along with the source code, is available at https://github.com/Nirmal2310/ANOMALY.