DeePhase
DeePhase predicts protein propensity for homotypic liquid-liquid phase separation (LLPS) from amino acid sequence to inform study of biomolecular condensates and cellular compartmentalization.
Key Features:
- In Silico Strategy: Performs global in silico analysis of associations between protein sequences and phase behavior.
- Machine Learning Models: Integrates machine learning approaches, including neural network-based language models, to predict LLPS propensity.
- Sequence Feature Analysis: Identifies LLPS-prone sequences as more intrinsically disordered, less hydrophobic, and lower in Shannon entropy compared with entries from the Protein Data Bank (PDB) and Swiss-Prot, and emphasizes a balance between polar and hydrophobic residues.
- Integrated Model: Combines knowledge-based sequence features with unsupervised embeddings from neural network models to distinguish LLPS-prone sequences from structured proteins and from unstructured proteins with lower LLPS propensity.
- High-Accuracy Prediction: Demonstrates high accuracy in identifying LLPS-prone sequences within the human proteome.
Scientific Applications:
- Cellular compartmentalization: Predicts proteins that may form biomolecular condensates to study mechanisms of cellular compartmentalization and regulation.
- Mechanistic hypothesis generation: Provides sequence-based hypotheses about molecular determinants of phase separation (disorder, hydrophobicity, Shannon entropy, polar/hydrophobic balance) for experimental testing.
- Proteome- and disease-focused screening: Enables proteome-scale identification of candidate phase-separating proteins to investigate roles in biological processes and diseases.
Methodology:
Trains classifiers on knowledge-based sequence features and unsupervised embeddings derived from neural network-based language models, performs global in silico association analyses between sequence and phase behavior, and validates predictions against established databases such as the Protein Data Bank and Swiss-Prot.
Topics
Details
- License:
- CC-BY-NC-4.0
- Tool Type:
- command-line tool, web application
- Programming Languages:
- Python
- Added:
- 9/8/2021
- Last Updated:
- 9/12/2021
Operations
Publications
Saar KL, Morgunov AS, Qi R, Arter WE, Krainer G, Lee AA, Knowles TPJ. Learning the molecular grammar of protein condensates from sequence determinants and embeddings. Proceedings of the National Academy of Sciences. 2021;118(15). doi:10.1073/pnas.2019053118. PMID:33827920. PMCID:PMC8053968.