SimGBS: a rapid method for simulating large-scale genotyping-by-sequencing data
Abstract
BACKGROUND: Genotyping-by-sequencing (GBS) offers a cost-effective solution to access genomic information on many samples. The advent of GBS has promoted large scale genomic analyses, including linkage analysis, heritability estimation, inference of genetic relatedness and genome-wide prediction and association studies, in many non-model species. Low-depth GBS can be used as a cost-effective method to obtain genotypes; however, not all alleles will be captured during the sequencing process, which can lead to high levels of genotype uncertainty. New methods that take genotype uncertainty into account are needed, but these methods can be difficult to develop and benchmark on real data. Simulations provide a convenient way to benchmark different approaches, and therefore, are essential for assessing existing methods and guiding future method development. RESULTS: SimGBS is a method for simulating large-scale GBS data from any real (e.g., reference) or simulated diploid genome. It is implemented in Julia but can also be accessed from R through the JuliaCall interface. Users can define populations with distinct demographic histories by modifying population size over a specified number of generations or providing a pedigree to generate structured populations. Most GBS parameters are highly customisable, including the choice of restriction enzyme and sequencing depth. SimGBS outputs both true genotype calls and DNA sequences in FASTQ format. The method is computationally efficient, capable of generating datasets for thousands of individuals; for example, simulating 172 samples required less than two hours while successfully reproducing the complexities observed in empirical GBS data. CONCLUSION: SimGBS is a valuable tool to assist with designing GBS experiments or evaluating bioinformatic pipelines and statistical methods for GBS data analysis in large scale population genetics studies.