eCOMPASS

eCOMPASS evaluates multiple sequence alignments (MSAs) of proteins using a direct coupling analysis (DCA)-based statistical measure to quantify alignment quality and detect biologically relevant residue–residue couplings reflected in 3D protein contacts.


Key Features:

  • DCA-based scoring: Uses direct coupling analysis (DCA) to generate scores for residue–residue couplings within MSAs.
  • Statistical significance testing: Calculates the statistical significance of congruence between high-scoring directly coupled pairs and corresponding 3D contacts in protein structures.
  • Alignment quality metrics: Quantifies alignment quality through DCA-derived metrics to distinguish true evolutionary couplings from random noise.
  • Large MSA applicability: Designed to assess large and precise MSAs typical in proteomics studies.
  • Validation datasets: Demonstrated on both simulated and real-world MSAs.
  • Implementation: Implemented in C++.

Scientific Applications:

  • MSA validation against structures: Assess whether MSAs align homologous residues consistent with 3D protein contacts.
  • Detection of coevolving residues: Identify compensatory substitutions and coevolving residue pairs indicative of structural or functional constraints.
  • Comparative evaluation of alignments: Compare the relative quality of alternative protein MSAs.
  • Large-scale proteomics analyses: Support construction and validation of large-scale MSAs to detect subtle biologically relevant patterns.

Methodology:

eCOMPASS employs direct coupling analysis (DCA) to score residue pairs, calculates the statistical significance of congruence between high-scoring coupled pairs and 3D protein contacts, and quantifies alignment quality using DCA-derived metrics; the software is implemented in C++.

Topics

Details

Programming Languages:
C++
Added:
9/8/2021
Last Updated:
11/24/2024

Operations

Publications

Neuwald AF, Kolaczkowski BD, Altschul SF. eCOMPASS: evaluative comparison of multiple protein alignments by statistical score. Bioinformatics. 2021;37(20):3456-3463. doi:10.1093/bioinformatics/btab374. PMID:33983436. PMCID:PMC8545322.

PMID: 33983436
PMCID: PMC8545322
Funding: - National Institute of General Medical Sciences: R01 GM125878 - National Science Foundation: BIO MCB 1817942

Downloads