SIMAP
SIMAP provides pre-calculated protein sequence similarities and associated annotations to enable large-scale protein similarity analysis and functional annotation.
Key Features:
- Pre-calculated sequence similarities: A comprehensive matrix of pre-computed pairwise protein similarities interconnecting the known protein sequence universe.
- Pre-computed protein features and domains: Protein features and domain annotations are provided for sequences across the database.
- Similarity clusters: Pre-defined similarity clusters group related protein sequences.
- Functional annotations: Functional annotations are provided, including those derived by the BLAST2GO algorithm.
- Coverage and scale: As of September 2013, the resource contains over 163 million proteins representing approximately 70 million non-redundant sequences.
- Data integration: Integrates data from all major public protein databases and consistently re-annotated metagenomes.
- Alignment methods: Employs sensitive FASTA search heuristics and Smith-Waterman alignments for similarity assessment.
- Domain models: Uses InterPro database models for protein domain annotation.
- All-against-all matrices: Ability to generate all-against-all protein sequence similarity matrices for specified protein collections.
- Protein similarity networks: Provides extraction of protein similarity networks for analysis of relationships and functions.
Scientific Applications:
- Protein sequence exploration: Comparative exploration of protein sequences across the global protein sequence universe.
- Functional annotation: Assignment and propagation of functional terms using BLAST2GO and InterPro-derived domain information.
- Comparative analyses: Large-scale all-against-all similarity analyses for project-specific protein collections.
- Network analysis: Construction and analysis of protein similarity networks to investigate protein relationships and functional modules.
- Non-redundant sequence mining: Analysis and mining of non-redundant sequence space at large scale.
Methodology:
Pre-calculation of sequence similarities using sensitive FASTA search heuristics and Smith-Waterman alignments; annotation with InterPro domain models and BLAST2GO functional annotation; generation of all-against-all similarity matrices, similarity clustering, and extraction of protein similarity networks.
Topics
Details
- Maturity:
- Emerging
- Cost:
- Free of charge
- Tool Type:
- web application
- Operating Systems:
- Linux, Windows, Mac
- Programming Languages:
- Java
- Added:
- 11/5/2015
- Last Updated:
- 1/10/2019
Operations
Publications
Arnold R, Goldenberg F, Mewes H, Rattei T. SIMAP—the database of all-against-all protein sequence similarities and annotations with new interfaces and increased coverage. Nucleic Acids Research. 2013;42(D1):D279-D284. doi:10.1093/nar/gkt970. PMID:24165881. PMCID:PMC3965014.