Panaroo

Panaroo constructs pangenomes using graph-based clustering to correct annotation and assembly errors and produce accurate comparative pangenome estimates for prokaryotic genomes.


Key Features:

  • Graph-Based Pangenome Clustering: Panaroo employs a graph-based clustering approach to cluster pangenomes and account for annotation errors introduced during genome assembly.
  • Error Mitigation: The method explicitly mitigates the effects of annotation and assembly errors and addresses complexities arising from horizontal gene transfer, gene duplication, and gene loss.
  • Extensive Validation: Panaroo has been validated using simulations of de novo assemblies under the infinitely many genes model and tested on multiple publicly available large bacterial genome datasets.
  • Negative Control Analysis (Mycobacterium tuberculosis): Use of a highly clonal Mycobacterium tuberculosis dataset demonstrates that uncorrected annotation errors can heavily bias pangenome estimates.
  • Enhanced Graphical Output: The tool produces graphical outputs that facilitate interpretation of pangenome structure and support downstream analyses such as pan-genome wide association studies and visualization of gene gain and loss rates.
  • PGWAS Support (Neisseria gonorrhoeae): Outputs have been used to perform pan-genome wide association studies, as demonstrated in Neisseria gonorrhoeae analyses.
  • Gene Gain and Loss Analysis (Pneumococcus): Panaroo has been applied to quantify gene gain and loss rates across 51 major global pneumococcal sequence clusters.

Scientific Applications:

  • Pan-Genome Wide Association Studies (PGWAS): Enables identification of genetic determinants associated with phenotypic traits across species, demonstrated for Neisseria gonorrhoeae.
  • Gene Gain and Loss Analysis: Supports quantification and comparison of gene gain and loss rates across bacterial sequence clusters, exemplified by analyses of 51 pneumococcal clusters.
  • Accurate Pangenome Estimation: Produces refined pangenome estimates for comparative genomics of prokaryotic genomes by correcting annotation and assembly artifacts.

Methodology:

Panaroo uses a graph-based clustering approach to distinguish true genetic variation from annotation artifacts and has been validated with simulations of de novo assemblies under the infinitely many genes model, tested on public large bacterial genome datasets, and evaluated using a clonal Mycobacterium tuberculosis negative control.

Topics

Details

License:
MIT
Programming Languages:
Python
Added:
1/18/2021
Last Updated:
1/22/2021

Operations

Publications

Tonkin-Hill G, MacAlasdair N, Ruis C, Weimann A, Horesh G, Lees JA, Gladstone RA, Lo S, Beaudoin C, Floto RA, Frost SD, Corander J, Bentley SD, Parkhill J. Producing Polished Prokaryotic Pangenomes with the Panaroo Pipeline. Unknown Journal. 2020. doi:10.1101/2020.01.28.922989.

Links