SEPATH
SEPATH detects and taxonomically classifies microbial sequences in human tissue whole-genome sequencing (WGS) data to enable pathogen identification in host-dominated metagenomic samples.
Key Features:
- High-Throughput Capability: Optimized for ultra-high-throughput analysis of large volumes of raw sequencing data from initial classification to final results.
- Python 3 Implementation and Snakemake Workflow: Implemented in Python 3 and managed with Snakemake for reproducible and scalable workflow execution.
- Benchmarked Performance: Evaluated across over 70 combinations of tools and parameters on 100 simulated cancer datasets spiked with realistic bacterial proportions, reporting top-performing tools and genus-level F1 scores.
- Top Tool Integration: Incorporates top-performing tools including mOTUs2 and Kraken, which achieved median genus-level F1 scores of 0.90 and 0.91 respectively for genus-level classification and bacterial proportion estimation.
- Pathogen-Relevant Detection: Targets detection of bacterial and viral sequences relevant to tumorigenesis, including Helicobacter pylori and human papillomavirus implicated in gastric non-cardia and cervical carcinomas.
- Post-Classification Filtering: Examines the impact of filtering parameters on result variability, particularly when using unassembled sequencing reads.
- HPC Optimization: Optimized for execution on high-performance computing clusters (e.g., the UEA cluster) for large-scale analyses.
- Repository of Scripts and Results: Includes a repository of scripts and results required for the computational analyses and evaluations described.
Scientific Applications:
- Cancer Metagenomics: Detection and taxonomic profiling of pathogens in tumor tissue WGS to investigate microbial associations with cancer and tumorigenesis.
- Host-Dominated Metagenomic Analysis: Identification and quantification of bacterial and viral sequences within host-dominated WGS datasets.
- Pathogen-Host Interaction Studies: Supporting analyses of pathogen presence and abundance to explore potential roles of microbes in disease mechanisms beyond cancer.
Methodology:
Implemented in Python 3 with Snakemake workflow management; evaluated via simulated benchmarking using 100 cancer datasets spiked with realistic bacterial proportions and over 70 tool/parameter combinations; integrates tools including mOTUs2 and Kraken; investigates post-classification filtering effects on unassembled sequencing reads and reports genus-level F1 performance metrics.
Topics
Details
- License:
- GPL-3.0
- Programming Languages:
- R, Shell, Python
- Added:
- 1/9/2020
- Last Updated:
- 12/19/2020
Operations
Publications
Gihawi A, Rallapalli G, Hurst R, Cooper CS, Leggett RM, Brewer DS. SEPATH: benchmarking the search for pathogens in human tissue whole genome sequence data leads to template pipelines. Genome Biology. 2019;20(1). doi:10.1186/s13059-019-1819-8. PMID:31639030. PMCID:PMC6805339.