MUFOLD

MUFOLD clusters de novo protein tertiary structure models to select near‑native representatives using novel distance measures Dscore1 and Dscore2.


Key Features:

  • Innovative distance measures: MUFOLD-CL computes Dscore1 and Dscore2 from protein distance matrices to assess model similarity in place of pairwise RMSD and TM-score calculations.
  • High correlation with established metrics: Dscore1 shows high correlation with RMSD and Dscore2 shows high correlation with TM-score, preserving accuracy in identifying near‑native models.
  • Linear time complexity: The Dscore1-based clustering algorithm scales linearly with the number of models, avoiding the quadratic time of pairwise comparisons.
  • Enhanced cluster representative selection: MUFOLD-CL uses Dscore2 to select cluster representatives, improving representative quality with minimal extra computation.
  • Fast data visualization: Dscore distribution–based visualization summarizes datasets of approximately 500,000 models in seconds to minutes.

Scientific Applications:

  • De novo model selection: Selecting near‑native structures from large populations generated by de novo protein structure prediction.
  • Structural biology analysis: Streamlining analysis of protein structure ensembles to study protein structure and dynamics.
  • Downstream research workflows: Providing representative models to support drug discovery, enzyme design, and investigations of disease mechanisms.

Methodology:

Compute Dscore1 and Dscore2 from protein distance matrices, replace traditional pairwise RMSD/TM-score calculations with matrix-based comparisons, apply Dscore1-based linear-time clustering, use Dscore2 to select cluster representatives, and generate Dscore distribution visualizations for large datasets.

Topics

Details

Tool Type:
command-line tool
Operating Systems:
Linux
Programming Languages:
Python
Added:
8/3/2017
Last Updated:
11/25/2024

Operations

Publications

Zhang J, Xu D. Fast algorithm for population‐based protein structural model analysis. PROTEOMICS. 2013;13(2):221-229. doi:10.1002/pmic.201200334. PMID:23184517. PMCID:PMC3641909.

Documentation

Links