In silico serotyping of E. Coli from short read data identifies limited novel o-loci but extensive diversity of O:H serotype combinations within and between pathogenic lineages

Danielle J. Ingle, Mary Valcanis, Alex Kuzevski, Marija Tauschek, Michael Inouye, Tim Stinear, Myron M. Levine, Roy M. Robins-Browne, Kathryn E. Holt

Research output: Contribution to journalArticleResearchpeer-review

98 Citations (Scopus)


The lipopolysaccharide (O) and flagellar (H) surface antigens of Escherichia coli are targets for serotyping that have traditionally been used to identify pathogenic lineages. These surface antigens are important for the survival of E. coli within mammalian hosts. However, traditional serotyping has several limitations, and public health reference laboratories are increasingly moving towards whole genome sequencing (WGS) to characterize bacterial isolates. Here we present a method to rapidly and accurately serotype E. coli isolates from raw, short read WGS data. Our approach bypasses the need for de novo genome assembly by directly screening WGS reads against a curated database of alleles linked to known and novel E. coli O-groups and H-types (the EcOH database) using the software package SRST2. We validated the approach by comparing in silico results for 197 enteropathogenic E. coli isolates with those obtained by serological phenotyping in an independent laboratory. We then demonstrated the utility of our method to characterize isolates in public health and clinical settings, and to explore the genetic diversity of >1500 E. coli genomes from multiple sources. Importantly, we showed that transfer of O- and H-antigen loci between E. coli chromosomal backbones is common, with little evidence of constraints by host or pathotype, suggesting that E. coli ’strain space’ may be virtually unlimited, even within specific pathotypes. Our findings show that serotyping is most useful when used in combination with strain genotyping to characterize microevolution events within an inferred population structure.

Original languageEnglish
Article number64
Number of pages14
JournalMicrobial Genomics
Issue number7
Publication statusPublished - 11 Jul 2016
Externally publishedYes


  • Diversity
  • E. coli
  • Genotype
  • Phenotype
  • Serotype

Cite this