eCOMPASS
eCOMPASS evaluates multiple sequence alignments (MSAs) of proteins using a direct coupling analysis (DCA)-based statistical measure to quantify alignment quality and detect biologically relevant residue–residue couplings reflected in 3D protein contacts.
Key Features:
- DCA-based scoring: Uses direct coupling analysis (DCA) to generate scores for residue–residue couplings within MSAs.
- Statistical significance testing: Calculates the statistical significance of congruence between high-scoring directly coupled pairs and corresponding 3D contacts in protein structures.
- Alignment quality metrics: Quantifies alignment quality through DCA-derived metrics to distinguish true evolutionary couplings from random noise.
- Large MSA applicability: Designed to assess large and precise MSAs typical in proteomics studies.
- Validation datasets: Demonstrated on both simulated and real-world MSAs.
- Implementation: Implemented in C++.
Scientific Applications:
- MSA validation against structures: Assess whether MSAs align homologous residues consistent with 3D protein contacts.
- Detection of coevolving residues: Identify compensatory substitutions and coevolving residue pairs indicative of structural or functional constraints.
- Comparative evaluation of alignments: Compare the relative quality of alternative protein MSAs.
- Large-scale proteomics analyses: Support construction and validation of large-scale MSAs to detect subtle biologically relevant patterns.
Methodology:
eCOMPASS employs direct coupling analysis (DCA) to score residue pairs, calculates the statistical significance of congruence between high-scoring coupled pairs and 3D protein contacts, and quantifies alignment quality using DCA-derived metrics; the software is implemented in C++.
Topics
Details
- Programming Languages:
- C++
- Added:
- 9/8/2021
- Last Updated:
- 11/24/2024
Operations
Publications
Neuwald AF, Kolaczkowski BD, Altschul SF. eCOMPASS: evaluative comparison of multiple protein alignments by statistical score. Bioinformatics. 2021;37(20):3456-3463. doi:10.1093/bioinformatics/btab374. PMID:33983436. PMCID:PMC8545322.