Browse Papers

Search the literature by keyword, topic, or author. Filter to articles with full text available.

3689 results

Full text 2026

Batch correction for large-scale mass spectrometry imaging experiments

Sparre AA, Jensen ON.

<h4>Summary</h4>We assess batch correction methods for MALDI mass spectrometry imaging experiments. ComBAT reduced batch-related technical variance, maintained biological variation, and improved the overall score by 19.4%.<h4>Availability and implementation</h4>Methods are available in R. comBAT is used …

Read PDF
Bioinformatics
Full text 2026

NEFFy: a versatile tool for computing the number of effective sequences

Haghani M, Bhattacharya D, Murali TM.

<h4>Motivation</h4>A Multiple Sequence Alignment (MSA) contains fundamental evolutionary information that is useful in the prediction of structure and function of proteins and nucleic acids. The "Number of Effective Sequences" (NEFF) quantifies the diversity of sequences …

Read PDF
Bioinformatics
Full text 2026

SNaQ.jl: Improved scalability for level-1 phylogenetic network inference

Kolbow N, Kong S, Chafin T, et al.

<h4>Motivation</h4>Phylogenetic networks represent complex biological scenarios that are overlooked in trees, such as hybridization and horizontal gene transfer. Although numerous methods have been developed for phylogenetic network inference, their scalability is severely limited by the …

Read PDF
Bioinformatics
Full text 2026

SEMPLR: an R package for transcription factor binding prediction

Kenney GE, Sherpa RN, Burgess JD, et al.

<h4>Summary</h4>SEMPLR is an R package that predicts transcription factor binding and variant effects using SNP Effect Matrices (SEMs), providing efficient, genome-wide scoring, enrichment testing, and visualization tools for comprehensive analysis of regulatory sequences.<h4>Availability</h4>Available on GitHub …

Read PDF
Bioinformatics
Full text 2026

Trimmomatic: a decade of feature-rich, high-performance NGS read preprocessing

Beier S, Bolger AM, Bolger ME, et al.

<h4>Motivation</h4>Trimmomatic is a widely adopted tool for preprocessing high-throughput sequencing data, particularly from Illumina platforms. Since its original publication in 2014, the volume and complexity of sequencing data have increased dramatically, necessitating continuous tool evolution.<h4>Results</h4>We …

Read PDF
Bioinformatics
Full text 2026

Addressing pandemic-wide systematic errors in the SARS-CoV-2 phylogeny

Hunt M, Hinrichs AS, Anderson D, et al.

The majority of SARS-CoV-2 genomes obtained during the pandemic were derived by amplifying overlapping windows of the genome ('tiled amplicons'), reconstructing their sequences and fitting them together. This leads to systematic errors in genomes unless …

Read PDF
Bioinformatics
Full text 2026

Vertebrate biodiversity via eDNA at the air-water interface

Ip YCA, Brandão-Dias PFP, Guri G, et al.

Aquatic, aerial, and terrestrial habitats exist along a continuum, with biomass and energy flows transporting genetic material across environmental boundaries. Here, we use environmental DNA (eDNA) metabarcoding to characterize genetic information exchange between water and …

Read PDF
Bioinformatics
Full text 2026

Efficient downsampling of genome alignments with Rasusa

Furqon ADC, Roberts LW, Hall MB.

High-throughput sequencing datasets frequently exhibit extreme read depth variation, biasing downstream analysis. Normalising coverage to a specific depth cap is important, yet existing tools rely on computationally expensive fetch-based or non-deterministic greedy algorithms. Here, we …

Read PDF
Bioinformatics Nanopore Sequencing