adabmDCA

adabmDCA infers maximum-entropy Potts and Ising statistical models from multiple sequence alignments (MSAs) to extract couplings and fields that characterize residue conservation and epistatic coevolution in protein and RNA families.


Key Features:

  • Model Inference: Estimates local fields (biases) and pairwise couplings for Potts/Ising variables from MSAs to capture residue conservation and epistasis.
  • Three-Dimensional Contact Map Prediction: Uses inferred couplings to predict residue–residue contacts and three-dimensional contact maps of target domains.
  • Mutational Effect Assessment and Sequence Generation: Evaluates models for predicting mutational effects and for generating in silico functional sequences.
  • Adaptive Learning Framework: Implements Boltzmann machine learning with both equilibrium and out-of-equilibrium learning methods to accommodate different computational constraints.
  • Parameter Pruning: Applies an information-based criterion to prune irrelevant parameters and reduce model complexity.

Scientific Applications:

  • Protein and RNA Families: Applied to MSAs of protein and RNA families to study conservation and coevolution.
  • Domain-Specific Modeling: Used to model the Kunitz and Beta-lactamase2 protein domains and the TPP-riboswitch RNA domain.
  • Structural and Evolutionary Studies: Employed for inferring contact maps and generating synthetic sequences relevant to structural biology, evolutionary analysis, and synthetic sequence design.

Methodology:

Boltzmann machine learning with gradient-ascent optimization of the likelihood, model observables computed via Markov Chain Monte Carlo (MCMC) sampling, support for equilibrium and out-of-equilibrium learning, and an information-based parameter pruning criterion.

Topics

Details

License:
MIT
Cost:
Free of charge
Tool Type:
command-line tool
Programming Languages:
C++
Added:
3/28/2022
Last Updated:
3/28/2022

Operations

Publications

Muntoni AP, Pagnani A, Weigt M, Zamponi F. adabmDCA: adaptive Boltzmann machine learning for biological sequences. BMC Bioinformatics. 2021;22(1). doi:10.1186/s12859-021-04441-9. PMID:34715775. PMCID:PMC8555268.

PMID: 34715775
PMCID: PMC8555268
Funding: - Simons Foundation: #454955 - Horizon 2020 Framework Programme: 734439 InferNet

Documentation

Links