SIMAP

SIMAP provides pre-calculated protein sequence similarities and associated annotations to enable large-scale protein similarity analysis and functional annotation.


Key Features:

  • Pre-calculated sequence similarities: A comprehensive matrix of pre-computed pairwise protein similarities interconnecting the known protein sequence universe.
  • Pre-computed protein features and domains: Protein features and domain annotations are provided for sequences across the database.
  • Similarity clusters: Pre-defined similarity clusters group related protein sequences.
  • Functional annotations: Functional annotations are provided, including those derived by the BLAST2GO algorithm.
  • Coverage and scale: As of September 2013, the resource contains over 163 million proteins representing approximately 70 million non-redundant sequences.
  • Data integration: Integrates data from all major public protein databases and consistently re-annotated metagenomes.
  • Alignment methods: Employs sensitive FASTA search heuristics and Smith-Waterman alignments for similarity assessment.
  • Domain models: Uses InterPro database models for protein domain annotation.
  • All-against-all matrices: Ability to generate all-against-all protein sequence similarity matrices for specified protein collections.
  • Protein similarity networks: Provides extraction of protein similarity networks for analysis of relationships and functions.

Scientific Applications:

  • Protein sequence exploration: Comparative exploration of protein sequences across the global protein sequence universe.
  • Functional annotation: Assignment and propagation of functional terms using BLAST2GO and InterPro-derived domain information.
  • Comparative analyses: Large-scale all-against-all similarity analyses for project-specific protein collections.
  • Network analysis: Construction and analysis of protein similarity networks to investigate protein relationships and functional modules.
  • Non-redundant sequence mining: Analysis and mining of non-redundant sequence space at large scale.

Methodology:

Pre-calculation of sequence similarities using sensitive FASTA search heuristics and Smith-Waterman alignments; annotation with InterPro domain models and BLAST2GO functional annotation; generation of all-against-all similarity matrices, similarity clustering, and extraction of protein similarity networks.

Topics

Details

Maturity:
Emerging
Cost:
Free of charge
Tool Type:
web application
Operating Systems:
Linux, Windows, Mac
Programming Languages:
Java
Added:
11/5/2015
Last Updated:
1/10/2019

Operations

Publications

Arnold R, Goldenberg F, Mewes H, Rattei T. SIMAP—the database of all-against-all protein sequence similarities and annotations with new interfaces and increased coverage. Nucleic Acids Research. 2013;42(D1):D279-D284. doi:10.1093/nar/gkt970. PMID:24165881. PMCID:PMC3965014.

Documentation