HapAsmbl: A reference-aided pipeline for assembling haplotypes in Nanopore amplicon sequence data of polymorphic populations
Abstract
<h4>Premise</h4>Advances in long-read sequencing offer new possibilities to investigate haplotype diversity across multiple genes in plants and other taxa through multi-locus, long-read amplicon sequencing (multi-locus LRAS). Despite this progress, there is a notable absence of dedicated bioinformatics pipelines for assembling diploid haplotypes of heterozygous individuals from such multi-locus LRAS datasets, which is required for highly polymorphic populations.<h4>Methods</h4>We first evaluated various de novo and reference-based assembly methods, culminating in a custom pipeline (HapAsmbl) to assemble haplotypes from Oxford Nanopore Technologies (ONT) LRAS data of five flowering genes (<i>FT3</i>, <i>FTL9</i>, <i>VRN1</i>, <i>VRN2A</i>, and <i>VRN2B</i>) generated from perennial ryegrass, a highly heterozygous species. After verifying the efficacy using a simulated heterozygous dataset, the HapAsmbl pipeline was used to explore haplotype diversity of <i>CO</i>, <i>FT3</i>, and <i>VRN1</i> across multiple ryegrass populations.<h4>Results</h4>HapAsmbl outperformed existing tools by reliably reconstructing diploid haplotypes across multiple loci, enabling efficient haplotype characterization and novel allele discovery in genetically diverse populations.<h4>Discussion</h4>HapAsmbl simplifies haplotype resolution from complex LRAS datasets from heterozygous individuals, allowing routine use of ONT long-read sequencing for scalable haplotype analysis. HapAsmbl will enable researchers to uncover novel alleles and relate these to phenotype, supporting plant-breeding efforts in non-model crops.