MAPGAPS

MAPGAPS performs global alignment of protein sequences using multiply-aligned profiles to detect, classify, and produce multiple sequence alignments of distantly related and functionally divergent proteins for evolutionary and functional analyses.


Key Features:

  • Statistical analysis of sequence patterns: Infers biochemical similarities and differences from patterns of sequence similarity and divergence among evolutionarily related proteins.
  • Use of multiply-aligned profiles: Employs multiply-aligned profiles as queries for detecting and classifying related sequences and as templates for multiple sequence alignment.
  • Karlin-Altschul statistics and PSI-BLAST heuristics: Applies Karlin-Altschul statistics to enhance sensitivity and incorporates PSI-BLAST heuristics to improve speed.
  • Handling of distantly related sequences: Detects and aligns weakly conserved, distantly related proteins, demonstrated by alignment of 33 P-loop GTPases with known structures.
  • Comparison with other methods: Outperformed sequence- and structure-based alignment methods such as hmmalign and PROMALS3D in comparative studies by avoiding misalignment of significant regions.
  • Scalability: Applied to a dataset of 65 million protein sequences to identify, classify, and align nearly half a million putative P-loop GTPase sequences.

Scientific Applications:

  • Phylogenetic analysis: Produces multiple sequence alignments suitable for reconstructing evolutionary relationships among protein families.
  • Functional annotation: Detects conserved and weakly conserved motifs to support inference of biochemical function and divergence.
  • Structural biology investigations: Aligns sequences with known structures, such as P-loop GTPases, to inform structural interpretation and comparative modeling.

Methodology:

Uses a curated multiple-profile alignment as input, employs multiply-aligned profiles as queries and templates, applies Karlin-Altschul statistics for scoring, and uses PSI-BLAST heuristics to accelerate searches while detecting weakly conserved sequence motifs across distantly related proteins.

Topics

Details

Tool Type:
command-line tool
Operating Systems:
Linux
Programming Languages:
C++
Added:
8/3/2017
Last Updated:
11/25/2024

Operations

Publications

Neuwald AF. Rapid detection, classification and accurate alignment of up to a million or more related protein sequences. Bioinformatics. 2009;25(15):1869-1875. doi:10.1093/bioinformatics/btp342. PMID:19505947. PMCID:PMC2732367.

Documentation

Links