Full text 2026

SquiDBase: a community resource of raw nanopore data from microbes

Cuypers WL, Ceylan H, Turcksin E, et al.

Full text

Loading PDF… Expand reader Download

Abstract

Nucleotide sequences in the FASTQ or BAM format are widely shared, yet derived from platform-specific raw data outputs that differ across sequencing platforms. In Oxford Nanopore Technologies (ONT) sequencing, raw signal data contain valuable biological information and enable basecaller optimization and modification detection. These raw signals also underpin algorithms that could improve ONT device portability and enhance target enrichment efficiency through adaptive sampling. Nevertheless, the storage and sharing of raw nanopore data remain limited due to technical constraints and the lack of standardized and centralized infrastructure. To address this challenge, we developed SquiDBase (https://squidbase.org), a dedicated repository for raw microbial nanopore sequencing data with linked processed data and metadata. To maximize immediate utility, we built SquiDPipe, a Nextflow pipeline for the automated removal of human reads from raw nanopore data, sequenced 24 clinically relevant viruses and incorporated them into SquiDBase, and added publicly available reference datasets and new community contributions. By offering a centralized, open-access raw data collection platform, SquiDBase facilitates data sharing, enhances reproducibility, and supports the development and benchmarking of computational tools, reinforcing open science in nanopore sequencing.