Cordax

Cordax maps amyloid-forming sequence space using machine learning and high-resolution structural data from the Protein Data Bank (PDB) to identify sequence determinants of amyloid propensity.


Key Features:

  • PDB-derived structural data: Leverages high-resolution amyloid core structures from the Protein Data Bank (PDB) for sequence analysis.
  • Machine learning: Applies machine learning approaches to explore amyloid-forming capabilities beyond established sequence biases.
  • t-SNE clustering: Clusters sequences using t-Distributed Stochastic Neighbour Embedding (t-SNE) for dimensionality reduction and visualization.
  • Sequence-space expansion: Identifies amyloid-capable sequences beyond traditional hydrophobic, beta-sheet-prone motifs and Q/N/Y-rich yeast prions.
  • Aliphatic and charge profiling: Detects clusters characterized by lower aliphatic content and higher net charge.
  • Secondary-structure tendencies: Identifies clusters with tendencies toward helical structures and disordered conformations.
  • Amyloid–solubility uncoupling: Reveals sequence contexts where amyloid propensity is uncoupled from solubility.
  • Surface-exposed compatibility: Highlights sequence motifs compatible with surface-exposed patches in globular proteins.
  • Functional and phase behavior links: Identifies motifs relevant to functional amyloids and sequences involved in liquid-liquid phase transitions.

Scientific Applications:

  • Structural biology of amyloids: Provides perspectives on sequence determinants and structural features of amyloid cores.
  • Protein misfolding diseases: Informs investigation of sequence variants and motifs relevant to amyloid-associated pathologies.
  • Functional amyloids: Aids identification of sequences that form biologically functional amyloids.
  • Liquid-liquid phase separation: Supports study of sequences involved in liquid-liquid phase transitions and related biophysical phenomena.
  • Sequence-space discovery: Expands the catalog of amyloid-compatible sequences for downstream experimental or computational study.

Methodology:

Uses high-resolution amyloid core structures from the Protein Data Bank (PDB) and applies machine learning, including clustering with t-Distributed Stochastic Neighbour Embedding (t-SNE) for dimensionality reduction and visualization.

Topics

Collections

Details

Added:
12/10/2020
Last Updated:
11/24/2024

Operations

Publications

Louros N, Orlando G, De Vleeschouwer M, Rousseau F, Schymkowitz J. Structure-based machine-guided mapping of amyloid sequence space reveals uncharted sequence clusters with higher solubilities. Nature Communications. 2020;11(1). doi:10.1038/s41467-020-17207-3. PMID:32620861. PMCID:PMC7335209.