UDSMProt

UDSMProt applies self-supervised deep sequence modeling to classify proteins from amino acid sequences and to learn universal sequence representations for downstream protein classification tasks.


Key Features:

  • Self-Supervised Pretraining: Employs a language-modeling self-supervised pretraining step on unlabeled protein sequences from Swiss-Prot to learn sequence representations.
  • Task-Agnostic Representation: Produces generalizable, task-agnostic protein sequence embeddings that can be reused across different classification tasks.
  • Fine-Tuning for Specific Tasks: Fine-tunes pretrained representations by updating model parameters to adapt to particular protein classification tasks.

Scientific Applications:

  • Enzyme Class Prediction: Classifies proteins into enzyme classes using only amino acid sequence information.
  • Gene Ontology Prediction: Predicts Gene Ontology terms covering biological processes, cellular components, and molecular functions from sequence data.
  • Remote Homology and Fold Detection: Detects remote homologies and protein folds to support inference of evolutionary relationships and structural similarity.

Methodology:

Pretraining: self-supervised language-modeling on unlabeled Swiss-Prot sequences; Fine-tuning: supervised adaptation of the pretrained model to specific classification tasks by updating model parameters.

Topics

Details

Programming Languages:
Python, Shell
Added:
1/18/2021
Last Updated:
3/6/2021

Operations

Publications

Strodthoff N, Wagner P, Wenzel M, Samek W. UDSMProt: universal deep sequence models for protein classification. Bioinformatics. 2020;36(8):2401-2409. doi:10.1093/bioinformatics/btaa003. PMID:31913448. PMCID:PMC7178389.

PMID: 31913448
PMCID: PMC7178389
Funding: - Berlin Big Data Center: 01IS14013A - Berlin Center for Machine Learning: 01IS18037I