SEPATH

SEPATH detects and taxonomically classifies microbial sequences in human tissue whole-genome sequencing (WGS) data to enable pathogen identification in host-dominated metagenomic samples.


Key Features:

  • High-Throughput Capability: Optimized for ultra-high-throughput analysis of large volumes of raw sequencing data from initial classification to final results.
  • Python 3 Implementation and Snakemake Workflow: Implemented in Python 3 and managed with Snakemake for reproducible and scalable workflow execution.
  • Benchmarked Performance: Evaluated across over 70 combinations of tools and parameters on 100 simulated cancer datasets spiked with realistic bacterial proportions, reporting top-performing tools and genus-level F1 scores.
  • Top Tool Integration: Incorporates top-performing tools including mOTUs2 and Kraken, which achieved median genus-level F1 scores of 0.90 and 0.91 respectively for genus-level classification and bacterial proportion estimation.
  • Pathogen-Relevant Detection: Targets detection of bacterial and viral sequences relevant to tumorigenesis, including Helicobacter pylori and human papillomavirus implicated in gastric non-cardia and cervical carcinomas.
  • Post-Classification Filtering: Examines the impact of filtering parameters on result variability, particularly when using unassembled sequencing reads.
  • HPC Optimization: Optimized for execution on high-performance computing clusters (e.g., the UEA cluster) for large-scale analyses.
  • Repository of Scripts and Results: Includes a repository of scripts and results required for the computational analyses and evaluations described.

Scientific Applications:

  • Cancer Metagenomics: Detection and taxonomic profiling of pathogens in tumor tissue WGS to investigate microbial associations with cancer and tumorigenesis.
  • Host-Dominated Metagenomic Analysis: Identification and quantification of bacterial and viral sequences within host-dominated WGS datasets.
  • Pathogen-Host Interaction Studies: Supporting analyses of pathogen presence and abundance to explore potential roles of microbes in disease mechanisms beyond cancer.

Methodology:

Implemented in Python 3 with Snakemake workflow management; evaluated via simulated benchmarking using 100 cancer datasets spiked with realistic bacterial proportions and over 70 tool/parameter combinations; integrates tools including mOTUs2 and Kraken; investigates post-classification filtering effects on unassembled sequencing reads and reports genus-level F1 performance metrics.

Topics

Details

License:
GPL-3.0
Programming Languages:
R, Shell, Python
Added:
1/9/2020
Last Updated:
12/19/2020

Operations

Publications

Gihawi A, Rallapalli G, Hurst R, Cooper CS, Leggett RM, Brewer DS. SEPATH: benchmarking the search for pathogens in human tissue whole genome sequence data leads to template pipelines. Genome Biology. 2019;20(1). doi:10.1186/s13059-019-1819-8. PMID:31639030. PMCID:PMC6805339.

PMID: 31639030
PMCID: PMC6805339
Funding: - Big C Cancer Charity: 16-09R

Links