sepp
sepp performs phylogenetic placement of metagenomic short reads by inserting query sequences into an existing alignment and phylogenetic tree of full-length sequences to elucidate evolutionary relationships.
Key Features:
- Phylogenetic placement: Inserts query sequences (metagenomic short reads) into an existing alignment and phylogenetic tree of full-length sequences.
- Per-query alignment estimation: Estimates an alignment for each query sequence relative to the full-length sequences' alignment prior to placement.
- Placement optimization: Determines the optimal phylogenetic tree location for each query sequence.
- Compatibility with existing methods: Builds upon HMMALIGN+EPA, HMMALIGN+pplacer, and PaPaRa+EPA workflows.
- SATé-derived dataset decomposition (SATé-boosting): Uses a dataset decomposition strategy from SATé with an iterative divide-and-conquer approach to co-estimate alignments and trees and to boost placement performance.
- Improved accuracy for large evolutionary diameter: Enhances the accuracy of HMMALIGN+pplacer for inputs with large evolutionary diameter.
- Accuracy–speed trade-offs: Provides more accurate placements for short sequences under challenging conditions and comparable accuracy with faster runtimes for simpler cases.
- Validation scope: Demonstrated accuracy and computational efficiency on biological and simulated data when provided with highly accurate full-length alignments and trees, with noted performance decline on very large or highly diverse full-length sequence sets.
Scientific Applications:
- Phylogenetic placement of metagenomic reads: Assigns metagenomic short reads to positions on a reference phylogeny for evolutionary interpretation.
- Inference of evolutionary relationships: Elucidates relationships between query sequences and known species using placement within a full-length-sequence tree and alignment.
- Method validation and benchmarking: Supports evaluation of placement algorithms using biological and simulated datasets.
Methodology:
Performs a two-step process of estimating a per-query alignment relative to the full-length alignment and then identifying the optimal placement on the phylogenetic tree; implements HMMALIGN, EPA, pplacer, PaPaRa approaches and a SATé-derived iterative divide-and-conquer dataset decomposition (SATé-boosting) to co-estimate alignments and trees.
Topics
Details
- License:
- GPL-3.0
- Cost:
- Free of charge
- Tool Type:
- command-line tool
- Programming Languages:
- Python, Java
- Added:
- 3/25/2022
- Last Updated:
- 3/25/2022
Operations
Publications
MIRARAB S, NGUYEN N, WARNOW T. SEPP: SATé-Enabled Phylogenetic Placement. Biocomputing 2012. 2011. doi:10.1142/9789814366496_0024. PMID:22174280.
PMID: 22174280