tax2tree

tax2tree maps taxonomic group names from reference taxonomies onto de novo phylogenetic trees to improve taxonomic resolution and annotation of bacterial and archaeal sequence datasets.


Key Features:

  • Taxonomy transfer methodology: Implements a "taxonomy-to-tree" approach that transfers group names from reference taxonomies such as Greengenes, NCBI, and cyanoDB onto the topology of de novo phylogenetic trees, enabling integration of large datasets (e.g., 408,315 sequences).
  • Explicit rank information: Incorporates explicit rank designations from the NCBI taxonomy by prefixing group names with rank information to standardize rank assignment.
  • Improved classification accuracy: Enhances taxonomic resolution, improving classification for 75% of sequences by one or more ranks relative to the original NCBI taxonomy.
  • Candidate phyla assessment: Evaluates NCBI-defined candidate phyla and provides recommendations to consolidate redundantly named groups, including 34 identified consolidations.
  • Pipeline components: Uses tree inference, jackknifing, and taxonomy transfer processes as core computational steps within the pipeline.

Scientific Applications:

  • Large-scale microbial sequencing projects: Supports taxonomic annotation in megasequencing efforts such as the Human Microbiome Project and the Earth Microbiome Project.
  • Marker gene and metagenomic surveys: Improves interpretation of marker gene and metagenomic surveys, particularly for under-classified environmental sequences.
  • Microbial ecology and evolution: Provides an improved taxonomy framework to facilitate studies of bacterial and archaeal ecology and evolutionary relationships.

Methodology:

Computational steps include de novo tree inference and jackknifing, followed by a taxonomy-to-tree transfer of group names from Greengenes, NCBI, and cyanoDB onto the tree topology, incorporation of NCBI rank prefixes into group names, and evaluation of candidate phyla with consolidation recommendations; the approach was demonstrated on a dataset of 408,315 sequences.

Topics

Details

Tool Type:
command-line tool
Operating Systems:
Linux, Windows, Mac
Programming Languages:
Python
Added:
12/18/2017
Last Updated:
11/25/2024

Operations

Data Inputs & Outputs

Publications

McDonald D, Price MN, Goodrich J, Nawrocki EP, DeSantis TZ, Probst A, Andersen GL, Knight R, Hugenholtz P. An improved Greengenes taxonomy with explicit ranks for ecological and evolutionary analyses of bacteria and archaea. The ISME Journal. 2011;6(3):610-618. doi:10.1038/ismej.2011.139. PMID:22134646. PMCID:PMC3280142.

Documentation

Links