DescFold
DescFold predicts protein fold types from amino acid sequences by integrating multiple sequence- and structure-based descriptors with machine learning for template-based fold recognition.
Key Features:
- Machine Learning Integration: Uses Support Vector Machines (SVMs) to integrate multiple sequence- and structure-based descriptors for fold recognition.
- Profile-Sequence-Alignment-Based Descriptor: Utilizes Psi-blast e-values and bit scores to capture sequence alignment information.
- Sequence-Profile-Alignment-Based Descriptor: Employs Rps-blast e-values and bit scores to enhance sensitivity to sequence variations.
- Secondary Structure Element Alignment (SSEA): Aligns secondary structure elements to capture structural motif information.
- PROSITE Functional Motifs: Incorporates PROSITE motif occurrences as predictive features.
- Profile-Profile Alignment Score (PPA): Generates profile-profile alignment scores using the COMPASS algorithm treating one profile as a template with known 3D structure.
- Structure-Seeded Profile: Creates structure-seeded profiles via TM-align structural alignment and refines template profiles by searching structural neighbors.
- Profile-Structural-Profile-Alignment-Based Descriptor (PSPA): Uses COMPASS to align profiles and structural information for improved discrimination.
- Training and Testing: Trained on 1,835 diverse proteins from SCOP 1.73 and tested on 1,866 proteins from SCOP 1.75.
- Performance Metrics: At a false positive rate below 5%, recognizes structural homologs at the fold level for approximately 46% of test proteins and was benchmarked against LiveBench targets and the Lindahl dataset.
Scientific Applications:
- Structural annotation: Predicts protein folds from sequences to support structural annotation of proteins.
- Functional inference: Provides structural predictions that aid inference of protein function and interactions.
- Evolutionary analysis: Detects fold-level homology to assess evolutionary relationships.
- Drug discovery and enzyme design: Supplies structure-based insights without experimental structures to support drug discovery and enzyme design efforts.
- Genome annotation: Assists annotation of newly sequenced proteins by assigning probable folds.
Methodology:
SVMs are trained to integrate descriptors computed from Psi-blast, Rps-blast, SSEA, PROSITE motif occurrences, COMPASS-derived PPA and PSPA scores, and TM-align–derived structure-seeded profiles (including searches of structural neighbors); training used SCOP 1.73 (1,835 proteins) and testing used SCOP 1.75 (1,866 proteins) with evaluation including false positive rate thresholds and benchmarks such as LiveBench and the Lindahl dataset.
Topics
Details
- Tool Type:
- command-line tool
- Operating Systems:
- Linux, Windows, Mac
- Programming Languages:
- Java, Perl
- Added:
- 12/18/2017
- Last Updated:
- 11/25/2024
Operations
Publications
Yan R, Si J, Wang C, Zhang Z. DescFold: A web server for protein fold recognition. BMC Bioinformatics. 2009;10(1). doi:10.1186/1471-2105-10-416. PMID:20003426. PMCID:PMC2803855.