Full text 2026

Systematic benchmarking of dorado basecalling models for RNA modification detection with highly multiplexed nanopore sequencing

Diensthuber G, Milenkovic I, Llovera L, et al.

Full text

Loading PDF… Expand reader Download

Abstract

Nanopore direct RNA sequencing holds promise for advancing our understanding of the epitranscriptome. Recently, Oxford Nanopore Technologies released basecalling models capable of detecting N6-methyladenosine (m6A), inosine (I), pseudouridine (Ψ), and 5-methylcytosine (m5C). However, their performance and cross-reactivity with other modifications remain largely unexplored. Here, we systematically benchmark four available modification-aware basecalling models by evaluating their per-read and per-site predictions across synthetic molecules and biological samples from diverse species. Models performed well on highly modified, balanced synthetic constructs (AUC = 0.93-0.97, PR-AUC = 0.84-0.91), but their performance dropped sharply on unbalanced datasets that reflect modification abundances in biological samples (PR-AUC: 0.04-0.09). Analysis of in vivo rRNA samples confirmed this limitation, with false-discovery rate ranging from 50% to 100%, even after filtering with modification-free controls. We identify two major sources of false positives: cross-reactivities with other modifications and current alterations at sites neighbouring a modified residue. Finally, we demonstrate that basecalling error- and current-based methods can accurately detect modifications, offering effective alternatives for modifications lacking dedicated models. Our results highlight the utility and limitations of modification-aware basecalling models for RNA modification detection, and underscore the importance of including control samples to mitigate false-positive predictions.