Scimm
Scimm performs unsupervised clustering of metagenomic reads from environmental DNA sequencing using interpolated Markov models to group sequences by species origin.
Key Features:
- Unsupervised learning: Operates without training data and clusters reads based on intrinsic sequence patterns rather than reference genomes.
- Interpolated Markov models: Uses interpolated Markov models to capture sequence dependencies at multiple scales for improved discrimination of origin.
- Species-level read clustering: Groups reads that originate from the same species to support downstream taxonomic binning.
- Database independence: Does not rely on pre-existing sequenced genomes from public databases for clustering.
- PHY SCIMM hybrid approach: Combines SCIMM with supervised techniques such as Phymm to enhance clustering accuracy when evolutionarily close training genomes are available.
- Improved accuracy: Demonstrates higher clustering accuracy compared to earlier methods in reported evaluations.
Scientific Applications:
- Metagenomic binning: Assigns environmental sequencing reads to species-level clusters for reconstruction of community composition.
- Novel microbe discovery: Facilitates analysis of samples containing predominantly novel microbial taxa by not depending on reference genomes.
- Hybrid supervised-enhanced clustering: Uses PHY SCIMM to improve clustering accuracy in samples with representatives from well-characterized genera.
Methodology:
Unsupervised clustering using interpolated Markov models to analyze sequence patterns and cluster reads by similarity; PHY SCIMM integrates SCIMM with supervised Phymm classification.
Topics
Details
- Tool Type:
- command-line tool
- Operating Systems:
- Linux
- Programming Languages:
- Python
- Added:
- 12/18/2017
- Last Updated:
- 12/10/2018
Operations
Publications
Kelley DR, Salzberg SL. Clustering metagenomic sequences with interpolated Markov models. BMC Bioinformatics. 2010;11(1). doi:10.1186/1471-2105-11-544. PMID:21044341. PMCID:PMC3098094.