PanGIA

PanGIA analyzes shotgun metagenomics data to detect pathogens and characterize microbial communities for biosurveillance, clinical, environmental, and forensic applications.


Key Features:

  • Pathogen detection and community profiling: Detects pathogenic microorganisms within complex samples and supports microbial community profiling and comparative metagenomics analyses.
  • Taxonomic identification with high precision: Uses BWA and Minimap2 to map reads to reference genomes, enabling taxonomy assignment down to the strain level.
  • Statistical confidence in taxonomic assignments: Estimates confidence using two independent approaches: integration of coverage and uniqueness data across reference genomes, and comparison of experimental samples to negative-control samples when available.
  • Quantitative summary metrics: Computes and reports confidence scores, normalized read abundance, reference genome coverage, depth-of-coverage, and RPKM for detected taxa.
  • Performance and scalability: Reports superior positive predictive value, sensitivity, and specificity relative to k-mer, read-mapping, and marker-gene based classifiers and processes datasets up to five million paired-end reads in under an hour on standard hardware.

Scientific Applications:

  • Biosurveillance: Detection of pathogenic organisms in environmental and surveillance samples to support threat monitoring.
  • Clinical diagnostics: Identification of pathogens in clinical shotgun metagenomics specimens for diagnostic interpretation.
  • Public health monitoring and outbreak response: Rapid pathogen detection and confidence estimation to inform outbreak investigations and public health actions.
  • Environmental and ecological studies: Comparative metagenomics and community profiling to study microbial ecology in environmental samples.
  • Forensic investigations: Precise microbial identification and strain-level assignment to support forensic analyses.

Methodology:

Reads are mapped to reference genomes using BWA/Minimap2; taxonomic confidence is estimated by integrating coverage and uniqueness metrics and by comparing experimental samples to negative-control samples when available; quantitative metrics computed include normalized read abundance, reference genome coverage, depth-of-coverage, and RPKM.

Topics

Details

Programming Languages:
Python, R
Added:
1/18/2021
Last Updated:
1/22/2021

Operations

Publications

Li P, Russell JA, Yarmosh D, Shteyman AG, Parker K, Wood H, Aspinwall J, Winegar R, Davenport K, Lo C, Bagnoli J, Davis P, Jacobs JL, Chain PS. PanGIA: A Metagenomics Analytical Framework for Routine Biosurveillance and Clinical Pathogen Detection. Unknown Journal. 2020. doi:10.1101/2020.04.20.051813.

Downloads

Links