EMIRGE

EMIRGE reconstructs near–full-length small subunit (SSU) ribosomal RNA genes, notably 16S rRNA, from short-read sequencing data and estimates relative abundances of the reconstructed sequences for microbial community profiling.


Key Features:

  • Full-length SSU (16S rRNA) reconstruction: Reconstructs near–full-length ribosomal gene sequences from short-read data.
  • Abundance estimation: Estimates sequence abundances for reconstructed SSU sequences within samples.
  • Expectation-Maximization iterative refinement: Adapts the Expectation-Maximization (EM) algorithm to iteratively refine abundance estimates and read-origin probabilities.
  • Probabilistic read assignment: Calculates the probability that each sequencing read originated from specific reference sequences.
  • Consensus sequence correction: Uses read-origin probabilities to correct consensus sequences for each reference SSU sequence each iteration.
  • Read mapping with Bowtie: Maps reads to reference sequences using Bowtie alignments, with design support for alternative mappers.
  • Convergence behavior: Repeats mapping and estimation until convergence, typically after 40–80 iterations.
  • Sensitivity and taxonomic resolution: Detects low-abundance organisms (as low as ~0.01% of a community) and resolves taxa including members of uncultivated phyla.

Scientific Applications:

  • Microbial community profiling: Reconstructs and quantifies SSU rRNA sequences to profile community composition from short-read datasets.
  • Diversity and rare taxa detection: Identifies rare and low-abundance organisms and captures diversity across cultivated and uncultivated phyla.
  • Environmental perturbation studies: Resolves shifts in community composition under perturbations, including transitions such as acetate amendment-driven iron reduction to sulfate reduction.
  • Comparative community analysis: Differentiates community compositions across environmental states and large datasets with reproducible results across technical replicates.
  • Population dynamics: Detects taxa-level responses to selective pressures and temporal changes within microbial communities.

Methodology:

EMIRGE takes short-read sequencing data, maps reads to reference sequences using Bowtie alignments, and applies an adapted Expectation-Maximization algorithm that estimates expected abundances of SSU sequences and probabilities that each read originated from particular references; it uses those probabilities to correct consensus sequences and repeats mapping and estimation until convergence, typically after 40–80 iterations.

Topics

Details

License:
GPL-3.0
Cost:
Free of charge
Tool Type:
command-line tool
Programming Languages:
Python
Added:
3/21/2022
Last Updated:
11/24/2024

Operations

Publications

Miller CS, Handley KM, Wrighton KC, Frischkorn KR, Thomas BC, Banfield JF. Short-Read Assembly of Full-Length 16S Amplicons Reveals Bacterial Diversity in Subsurface Sediments. PLoS ONE. 2013;8(2):e56018. doi:10.1371/journal.pone.0056018. PMID:23405248. PMCID:PMC3566076.

Links