pydca

pydca infers coevolutionary couplings from multiple sequence alignments to identify residue–residue interactions for structural and functional analysis.


Key Features:

  • Algorithmic approaches: Implements inverse statistical methods—mean-field approximation and pseudo-likelihood maximization—to infer coevolutionary relationships between residues from MSAs.
  • Input formats: Accepts multiple sequence alignment (MSA) files in FASTA format.
  • Reference sequence mapping: Optionally maps a supplied reference sequence onto the MSA to compute coevolutionary scores for pairs of sites in the reference.
  • Preprocessing and visualization: Provides MSA trimming and contact map visualization to process alignments and display inferred residue contacts.

Scientific Applications:

  • Coevolutionary coupling inference: Identifies direct residue couplings indicative of evolutionary constraints and potential physical contacts.
  • Contact and structure prediction: Supports prediction of protein and RNA contacts useful for tertiary structure modeling.
  • Interpretation of sequence variation: Aids analysis of functional implications of sequence variants through changes in coupling patterns.
  • Support for experimental structure determination: Provides contact information that can guide experimental efforts such as mutagenesis and structure solving.

Methodology:

Applies mean-field approximation and pseudo-likelihood maximization to MSAs in FASTA format, with optional reference-sequence mapping, and includes MSA trimming and contact map visualization.

Topics

Details

License:
MIT
Programming Languages:
C++, Python, C
Added:
1/9/2020
Last Updated:
12/10/2020

Operations

Publications

Zerihun MB, Pucci F, Peter EK, Schug A. pydca v1.0: a comprehensive software for Direct Coupling Analysis of RNA and Protein Sequences. Unknown Journal. 2019. doi:10.1101/805523.