Full text 2026

A dual approach to evaluate the performance of RNA-Seq data analysis pipelines with weak signals

Baroudi M, Ly FO, Goujon E, et al.

Full text

Loading PDF… Expand reader Download

Abstract

<h4>Purpose</h4>In this study, we evaluated 90 bioinformatics pipelines using RNA-Seq datasets from Rats, Zebrafish and Mice. The analysis was conducted in the context of weak signals, including exposure to metallic particles (tungsten), low-dose radiation or medical treatment. RNA-Seq data analysis involves several critical steps, from quality control to differential expression analysis, each offering multiple algorithmic options. Selecting the optimal pipeline is particularly challenging in complex scenarios with weak signals.We applied a dual strategy based on two complementary approaches to rank and evaluate the performance of these pipelines. The first approach, a widely used method, is based on the correlation between RNA-Seq and qRT-PCR expression data to ensure the direct validation of RNA-Seq results. The second approach leverages machine learning classifiers to rank the pipelines based on their ability to distinguish between exposure groups. This dual strategy was designed to identify the most reliable pipelines capable of providing accurate biological insights, with the top-performing pipeline highlighting key biological processes linked to a weak signal.<h4>Results</h4>Our results highlight the crucial role of pipeline selection in RNA-Seq studies, as it influences both analysis efficiency and biological insights. While our findings are particularly relevant for studies with weak RNA-Seq signals, the ranking methods we employed can be applied in other fields to identify the most appropriate pipeline for generating biologically meaningful data. We provide practical recommendations for bioinformaticians to select robust pipelines, ensuring reliable and insightful outcomes across various research contexts, including environmental exposures.<h4>Conclusions</h4>RNA-Seq pipelines that are effective for strong signals may fail with weak signals or noisy RNA-seq data. Our classifier-based ranking approach is still particularly useful, even when the sequencing depth is lower. Pipelines' sensitivity has a greater impact on counting and normalization than on trimming and mapping. StringTie should be prioritized as a counting method for data with weak signals.