RFQAmodel
RFQAmodel assesses protein structure prediction model quality to identify correct folds among models generated by template-free prediction protocols by applying a Random Forest classifier that integrates existing quality assessment scores and two predicted contact map alignment scores.
Key Features:
- Quality Assessment Framework: A Random Forest-based classifier integrates existing quality assessment scores with two predicted contact map alignment scores to identify correct models, including cases where traditional methods underperform.
- Training and Testing: The classifier was trained on a large, structurally diverse and balanced dataset of protein domains and validated on an independent test set of 244 protein domains.
- Performance Metrics: When ranking models by RFQAmodel scores, the highest-ranking model had the correct fold for 67 targets (52 confirmed accurate), and the method correctly identified incorrect models for 59 targets.
- Comparative Advantage: The method excels at pinpointing correct models when only a few among many are accurate and showed consistent performance on CASP12 and CASP13 free-modelling model sets.
- Iterative Modeling Protocol: Supports an iterative generate-and-evaluate approach in which models are repeatedly generated and assessed with RFQAmodel until a high-confidence correct model is identified.
Scientific Applications:
- De novo structure prediction model selection: Prioritizes and ranks template-free (de novo) protein models to identify correct folds for downstream analysis.
- Resource prioritization for challenging targets: Enables iterative focusing of computational effort on difficult targets by identifying when additional modeling is warranted.
- Structural biology and structural genomics: Improves reliability of predicted structures used in structural biology studies and structural genomics projects.
Methodology:
RFQAmodel uses a Random Forest classifier that integrates existing quality assessment scores with two predicted contact map alignment scores; it was trained on a large, structurally diverse, balanced dataset of protein domains, validated on an independent test set of 244 domains, and applied to rank models (including CASP12 and CASP13 free-modelling sets) with support for iterative generate-and-evaluate cycles.
Topics
Details
- Added:
- 1/9/2020
- Last Updated:
- 1/15/2021
Operations
Publications
West CE, de Oliveira SHP, Deane CM. RFQAmodel: Random Forest Quality Assessment to identify a predicted protein structure in the correct fold. PLOS ONE. 2019;14(10):e0218149. doi:10.1371/journal.pone.0218149. PMID:31634369. PMCID:PMC6802825.