A dataset of 352 nuclear genes for accurate species identification and geographical origin traceability of <i>Rhododendron dauricum</i> L
Abstract
This dataset presents 352 nuclear genes assembled from whole genome skimming data of 43 <i>Rhododendron</i> samples. The data were generated from 14 <i>Rhododendron dauricum</i> collected from seven distinct geographical populations in Northeast China, together with sequence data from 29 additional <i>Rhododendron</i> samples downloaded from the NCBI database. Using the universal set of 353 angiosperm nuclear genes as a reference, all genes were assembled with the HybPiper v2.1.1 pipeline. The dataset contains raw assembly sequences in FASTA format for each gene. Sequence alignment, trimming, and phylogenetic analysis were performed to construct phylogenetic trees. The resulting phylogenies based on concatenated 352-gene dataset and the screened 17-gene sub-dataset clearly distinguished <i>R. dauricum</i> from other <i>Rhododendron</i> species. Moreover, both datasets resolved individuals from the same population into distinct clades, enabling geographical origin traceability for the protected species <i>R. dauricum</i>. This dataset provides high-resolution molecular markers for research on <i>Rhododendron</i> phylogenomics, population genetics, conservation, and molecular identification.