EMIRGE
EMIRGE reconstructs near–full-length small subunit (SSU) ribosomal RNA genes, notably 16S rRNA, from short-read sequencing data and estimates relative abundances of the reconstructed sequences for microbial community profiling.
Key Features:
- Full-length SSU (16S rRNA) reconstruction: Reconstructs near–full-length ribosomal gene sequences from short-read data.
- Abundance estimation: Estimates sequence abundances for reconstructed SSU sequences within samples.
- Expectation-Maximization iterative refinement: Adapts the Expectation-Maximization (EM) algorithm to iteratively refine abundance estimates and read-origin probabilities.
- Probabilistic read assignment: Calculates the probability that each sequencing read originated from specific reference sequences.
- Consensus sequence correction: Uses read-origin probabilities to correct consensus sequences for each reference SSU sequence each iteration.
- Read mapping with Bowtie: Maps reads to reference sequences using Bowtie alignments, with design support for alternative mappers.
- Convergence behavior: Repeats mapping and estimation until convergence, typically after 40–80 iterations.
- Sensitivity and taxonomic resolution: Detects low-abundance organisms (as low as ~0.01% of a community) and resolves taxa including members of uncultivated phyla.
Scientific Applications:
- Microbial community profiling: Reconstructs and quantifies SSU rRNA sequences to profile community composition from short-read datasets.
- Diversity and rare taxa detection: Identifies rare and low-abundance organisms and captures diversity across cultivated and uncultivated phyla.
- Environmental perturbation studies: Resolves shifts in community composition under perturbations, including transitions such as acetate amendment-driven iron reduction to sulfate reduction.
- Comparative community analysis: Differentiates community compositions across environmental states and large datasets with reproducible results across technical replicates.
- Population dynamics: Detects taxa-level responses to selective pressures and temporal changes within microbial communities.
Methodology:
EMIRGE takes short-read sequencing data, maps reads to reference sequences using Bowtie alignments, and applies an adapted Expectation-Maximization algorithm that estimates expected abundances of SSU sequences and probabilities that each read originated from particular references; it uses those probabilities to correct consensus sequences and repeats mapping and estimation until convergence, typically after 40–80 iterations.
Topics
Details
- License:
- GPL-3.0
- Cost:
- Free of charge
- Tool Type:
- command-line tool
- Programming Languages:
- Python
- Added:
- 3/21/2022
- Last Updated:
- 11/24/2024
Operations
Publications
Miller CS, Handley KM, Wrighton KC, Frischkorn KR, Thomas BC, Banfield JF. Short-Read Assembly of Full-Length 16S Amplicons Reveals Bacterial Diversity in Subsurface Sediments. PLoS ONE. 2013;8(2):e56018. doi:10.1371/journal.pone.0056018. PMID:23405248. PMCID:PMC3566076.