Full text 2026

Development and comparison of evaluation metrics for batch correction reveals performance differences

Laiho A, Laitinen M, Holm L, et al.

Full text

Loading PDF… Expand reader Download

Abstract

<h4>Motivation</h4>Batch effects are a common challenge in the analysis of biological datasets, particularly RNA-seq data. Although numerous methods exist to correct batch effects, their comparative evaluation has received limited attention. Several metrics have been proposed to assess the effectiveness of batch correction, but it is unclear how consistently these metrics reflect performance. Here, we systematically investigate differences in the behavior and sensitivity of commonly used evaluation metrics for batch effect removal.<h4>Results</h4>We compiled a set of established evaluation metrics and introduced several new metrics. These were systematically compared across multiple datasets generated using our Artificial Dilution Series approach: each dataset contained controlled levels of noise simulating batch effects, enabling quantitative assessment of each metric's ability to discriminate between noise levels. We observed consistent differences among metrics, with those based on the F-statistic and the Davies-Bouldin index showing the strongest discriminative performance.<h4>Availability and implementation</h4>The data and codes used in this article are available in https://github.com/Aleksi95/BatchMetrics.