Sequedex
Sequedex classifies short reads from shotgun metagenomics by exact matching to a precomputed set of peptide 10-mers to provide phylogenetic and functional assignments.
Key Features:
- High-Speed Processing: Processes approximately 6.6 gigabase pairs per hour on a single CPU compared to about 25 kilobase pairs per hour for BLASTX against the NR database.
- Exact Matching Methodology: Uses exact matches between reads and peptide 10-mers, with about 20 million 10-mer peptides identified across 403 representative bacterial genomes.
- Phylogenetic Classification: Assigns each peptide 10-mer as a signature to nodes in a phylogenetic reference tree (primarily based on RNA polymerase genes) and classifies fragments by the most specific node consistent with observed signature peptides.
- Sensitivity and Specificity: Exhibits sensitivity and specificity comparable to MEGAN when using BLASTX against NR, validated on synthetic data from newly sequenced soil-bacterium genomes and ten real soil metagenomics datasets.
- Signal-to-Noise Ratio: Achieves a signal-to-noise ratio of approximately 400 in real metagenomics data, enabling discrimination among environmental samples.
- Versatility with Read Lengths: Produces results comparable to existing techniques for reads longer than ~300 base pairs while maintaining robust performance with shorter reads.
Scientific Applications:
- Metagenomic classification in high-diversity environments: Classifies short reads from environments with high sequence diversity and few close reference sequences, such as soils.
- Large-scale environmental genomics: Enables rapid assignment of reads for inclusion in high-throughput environmental and ecological genomic analysis pipelines.
- Phylogenetic and functional genomics studies: Provides peptide-based signature assignments that support phylogenetic placement and functional inference in microbial community analyses.
Methodology:
Precomputes peptide 10-mers from representative bacterial genomes (~403 genomes, ~20 million peptides), assigns each 10-mer as a signature to nodes of a phylogenetic reference tree primarily based on RNA polymerase genes, and classifies reads by exact matching translated fragments to those peptide 10-mers to identify the most specific consistent phylogenetic node.
Topics
Details
- Maturity:
- Mature
- Tool Type:
- desktop application
- Operating Systems:
- Linux, Mac
- Programming Languages:
- Java, Python
- Added:
- 1/13/2017
- Last Updated:
- 11/25/2024
Operations
Publications
Berendzen J, Bruno WJ, Cohn JD, Hengartner NW, Kuske CR, McMahon BH, Wolinsky MA, Xie G. Rapid phylogenetic and functional classification of short genomic fragments with signature peptides. BMC Research Notes. 2012;5(1). doi:10.1186/1756-0500-5-460. PMID:22925230. PMCID:PMC3772700.