MUFOLD_CL
MUFOLD_CL clusters protein structure models using two distance-matrix-derived metrics, Dscore1 and Dscore2, to enable efficient clustering and selection of near-native structures.
Key Features:
- Innovative Distance Measures: Dscore1 and Dscore2 are derived from protein distance matrices, with Dscore1 evaluating model differences and Dscore2 evaluating model similarities.
- High Correlation with Established Metrics: Dscore1 shows strong correlation with RMSD and Dscore2 shows strong correlation with TM-score, maintaining consistency with established metrics.
- Linear Time Complexity: Utilizing Dscore1 for clustering reduces algorithmic complexity to linear time with respect to the number of models.
- Enhanced Cluster Representation Quality: Dscore2 is used to select representative models from clusters, improving representative quality with minimal additional computation.
- Fast Data Visualization: Dscore distributions enable rapid visualization of very large model sets (e.g., ~500,000 models) to assess clustering outcomes.
Scientific Applications:
- Structural model management and quality assessment: Supports management and interpretation of large collections of protein structure models generated by de novo prediction methods and aids identification of near-native structures for studies of protein function and therapeutic design.
Methodology:
Computes Dscore1 and Dscore2 from protein distance matrices and compares models via these measures instead of traditional pairwise RMSD/TM-score calculations to reduce computational demands.
Topics
Details
- Tool Type:
- command-line tool
- Operating Systems:
- Linux
- Programming Languages:
- PHP
- Added:
- 8/3/2017
- Last Updated:
- 11/25/2024
Operations
Publications
Zhang J, Xu D. Fast algorithm for population‐based protein structural model analysis. PROTEOMICS. 2013;13(2):221-229. doi:10.1002/pmic.201200334. PMID:23184517. PMCID:PMC3641909.