sepp

sepp performs phylogenetic placement of metagenomic short reads by inserting query sequences into an existing alignment and phylogenetic tree of full-length sequences to elucidate evolutionary relationships.


Key Features:

  • Phylogenetic placement: Inserts query sequences (metagenomic short reads) into an existing alignment and phylogenetic tree of full-length sequences.
  • Per-query alignment estimation: Estimates an alignment for each query sequence relative to the full-length sequences' alignment prior to placement.
  • Placement optimization: Determines the optimal phylogenetic tree location for each query sequence.
  • Compatibility with existing methods: Builds upon HMMALIGN+EPA, HMMALIGN+pplacer, and PaPaRa+EPA workflows.
  • SATé-derived dataset decomposition (SATé-boosting): Uses a dataset decomposition strategy from SATé with an iterative divide-and-conquer approach to co-estimate alignments and trees and to boost placement performance.
  • Improved accuracy for large evolutionary diameter: Enhances the accuracy of HMMALIGN+pplacer for inputs with large evolutionary diameter.
  • Accuracy–speed trade-offs: Provides more accurate placements for short sequences under challenging conditions and comparable accuracy with faster runtimes for simpler cases.
  • Validation scope: Demonstrated accuracy and computational efficiency on biological and simulated data when provided with highly accurate full-length alignments and trees, with noted performance decline on very large or highly diverse full-length sequence sets.

Scientific Applications:

  • Phylogenetic placement of metagenomic reads: Assigns metagenomic short reads to positions on a reference phylogeny for evolutionary interpretation.
  • Inference of evolutionary relationships: Elucidates relationships between query sequences and known species using placement within a full-length-sequence tree and alignment.
  • Method validation and benchmarking: Supports evaluation of placement algorithms using biological and simulated datasets.

Methodology:

Performs a two-step process of estimating a per-query alignment relative to the full-length alignment and then identifying the optimal placement on the phylogenetic tree; implements HMMALIGN, EPA, pplacer, PaPaRa approaches and a SATé-derived iterative divide-and-conquer dataset decomposition (SATé-boosting) to co-estimate alignments and trees.

Topics

Details

License:
GPL-3.0
Cost:
Free of charge
Tool Type:
command-line tool
Programming Languages:
Python, Java
Added:
3/25/2022
Last Updated:
3/25/2022

Operations

Publications

MIRARAB S, NGUYEN N, WARNOW T. SEPP: SATé-Enabled Phylogenetic Placement. Biocomputing 2012. 2011. doi:10.1142/9789814366496_0024. PMID:22174280.

Documentation

Links