Prediction of genetic relatedness of <i>Escherichia coli</i> using neighbor typing: a tool for rapid outbreak detection
Abstract
Identifying the genetic relatedness of resistant bacterial pathogens in healthcare settings can help identify undetected transmission events and outbreaks. However, current methods are time- and resource-intensive. We evaluated a rapid neighbor typing method paired with long-read sequencing for assessment of genetic relatedness. Utilizing a data set of primary clinical samples and published isolate data from two outbreaks of <i>Escherichia coli</i>, we applied genomic neighbor typing of long-read sequence data to rapidly estimate genetic relatedness. We assessed the correlation between neighbor typing predicted genetic distance and pairwise genetic distance from short-read draft whole genomes for all sample pairs. Predicted genetic trees using neighbor typing were compared to reference genetic trees generated using mash distances and maximum-likelihood (ML) methods to assess the extent of agreement, along with metrics of cluster similarity (cluster comparability and Baker's gamma index [BGI]) and tree topology similarity (generalized Robinson-Foulds [GRF] metric). For all three data sets, we found strong correlations between the reference methods and predicted genetic distances (Spearman's rho = 0.75-0.95, <i>P</i> < 0.001), which improved when using a lineage score-informed approach (Spearman's rho = 0.93-0.94, <i>P</i> < 0.001). Predicted genetic trees and clusters from neighbor typing were comparable to those generated using either <i>mashtree</i> or an ML method, with a range of cluster comparability of 85.8-99.5%, BGIs of 0.8-0.95, and GRF values of 0.34-0.8. Pairing the neighbor typing method with long-read sequencing can enable accurate predictions of the relatedness of <i>E. coli</i> samples and isolates, and could potentially be used as a rapid outbreak surveillance tool.