EPA-ng

EPA-ng performs scalable phylogenetic placement of molecular sequences onto a reference phylogenetic tree to support taxonomic and evolutionary analysis of large next-generation sequencing datasets.


Key Features:

  • Evolutionary Placement Algorithm Implementation: Reimplements the evolutionary placement algorithm (EPA) for phylogenetic placement of query sequences onto a fixed reference tree.
  • High-Performance Phylogenetic Placement: Achieves substantial speed improvements compared with RAxML-EPA and pplacer for large-scale sequence datasets.
  • Distributed Memory Parallelization: Supports distributed memory parallelization for execution on high-performance computing clusters.
  • Shared and Distributed Memory Support: Executes on both shared-memory systems and distributed-memory computing environments.
  • Integration of EPA Methodologies: Incorporates algorithmic concepts from RAxML-EPA and pplacer to improve scalability and performance.

Scientific Applications:

  • Metagenomic Taxonomic Placement: Assigns environmental sequencing reads to positions on reference phylogenetic trees.
  • Microbial Diversity Analysis: Investigates evolutionary relationships within complex microbial communities.
  • Large-Scale Phylogenetic Studies: Processes large next-generation sequencing datasets for evolutionary and ecological analyses.

Methodology:

EPA-ng places query sequences onto a fixed reference phylogenetic tree using a reimplementation of the evolutionary placement algorithm with distributed and shared memory parallelization.

Topics

Details

License:
AGPL-3.0
Maturity:
Mature
Cost:
Free of charge
Tool Type:
command-line tool
Operating Systems:
Linux, Mac
Added:
5/31/2019
Last Updated:
6/16/2020

Operations

Publications

Barbera P, Kozlov AM, Czech L, Morel B, Darriba D, Flouri T, Stamatakis A. EPA-ng: Massively Parallel Evolutionary Placement of Genetic Sequences. Systematic Biology. 2018;68(2):365-369. doi:10.1093/sysbio/syy054. PMID:30165689. PMCID:PMC6368480.