APPLES

APPLES performs phylogenetic placement of DNA and protein sequences into pre-existing reference trees using kmer-based distances and least-squares optimization to enable scalable placement on large reference trees.


Key Features:

  • Scalability: Handles reference trees with a very large number of leaves, tested up to 200,000.
  • Speed and Efficiency: Achieves runtimes and memory usage an order of magnitude lower than traditional maximum likelihood (ML) methods on large datasets.
  • Non-dependency on sequence alignment: Does not require aligned query sequences or assembled reference sequences, enabling placement from unassembled reads.
  • kmer-based distance calculation: Uses kmer-based distances to compute pairwise sequence distances for placement without alignment.
  • Accuracy with dense reference trees: Demonstrates improved placement accuracy when using dense reference trees compared to ML methods on sparser trees.

Scientific Applications:

  • Phylogenetic updates: Scalable integration of new species data into existing phylogenies for taxonomy and tree updates.
  • Sample identification: Identification of unknown query samples using DNA and protein data from (meta-)barcoding and other sequencing sources.
  • Genome skimming applications: Placement of unassembled reads from genome skimming projects for sample identification.
  • Metagenomic analyses: Placement of sequences from metagenomic datasets for community profiling and sample characterization.

Methodology:

APPLES performs distance-based phylogenetic placement using kmer-based distance calculations and least-squares optimization without requiring sequence alignment.

Topics

Details

License:
MIT
Programming Languages:
Python
Added:
11/14/2019
Last Updated:
11/24/2024

Operations

Data Inputs & Outputs

Phylogenetic tree generation (maximum likelihood and Bayesian methods)

Publications

Balaban M, Sarmashghi S, Mirarab S. APPLES: Scalable Distance-Based Phylogenetic Placement with or without Alignments. Systematic Biology. 2019;69(3):566-578. doi:10.1093/sysbio/syz063. PMID:31545363. PMCID:PMC7164367.

PMID: 31545363
PMCID: PMC7164367
Funding: - National Science Foundation: IIS-1565862 - National Institutes of Health: 5P30AI027767-28 - NSF: ACI-1053575, NSF-1815485