Cordax
Cordax maps amyloid-forming sequence space using machine learning and high-resolution structural data from the Protein Data Bank (PDB) to identify sequence determinants of amyloid propensity.
Key Features:
- PDB-derived structural data: Leverages high-resolution amyloid core structures from the Protein Data Bank (PDB) for sequence analysis.
- Machine learning: Applies machine learning approaches to explore amyloid-forming capabilities beyond established sequence biases.
- t-SNE clustering: Clusters sequences using t-Distributed Stochastic Neighbour Embedding (t-SNE) for dimensionality reduction and visualization.
- Sequence-space expansion: Identifies amyloid-capable sequences beyond traditional hydrophobic, beta-sheet-prone motifs and Q/N/Y-rich yeast prions.
- Aliphatic and charge profiling: Detects clusters characterized by lower aliphatic content and higher net charge.
- Secondary-structure tendencies: Identifies clusters with tendencies toward helical structures and disordered conformations.
- Amyloid–solubility uncoupling: Reveals sequence contexts where amyloid propensity is uncoupled from solubility.
- Surface-exposed compatibility: Highlights sequence motifs compatible with surface-exposed patches in globular proteins.
- Functional and phase behavior links: Identifies motifs relevant to functional amyloids and sequences involved in liquid-liquid phase transitions.
Scientific Applications:
- Structural biology of amyloids: Provides perspectives on sequence determinants and structural features of amyloid cores.
- Protein misfolding diseases: Informs investigation of sequence variants and motifs relevant to amyloid-associated pathologies.
- Functional amyloids: Aids identification of sequences that form biologically functional amyloids.
- Liquid-liquid phase separation: Supports study of sequences involved in liquid-liquid phase transitions and related biophysical phenomena.
- Sequence-space discovery: Expands the catalog of amyloid-compatible sequences for downstream experimental or computational study.
Methodology:
Uses high-resolution amyloid core structures from the Protein Data Bank (PDB) and applies machine learning, including clustering with t-Distributed Stochastic Neighbour Embedding (t-SNE) for dimensionality reduction and visualization.
Topics
Collections
Details
- Added:
- 12/10/2020
- Last Updated:
- 11/24/2024
Operations
Publications
Louros N, Orlando G, De Vleeschouwer M, Rousseau F, Schymkowitz J. Structure-based machine-guided mapping of amyloid sequence space reveals uncharted sequence clusters with higher solubilities. Nature Communications. 2020;11(1). doi:10.1038/s41467-020-17207-3. PMID:32620861. PMCID:PMC7335209.