UDSMProt
UDSMProt applies self-supervised deep sequence modeling to classify proteins from amino acid sequences and to learn universal sequence representations for downstream protein classification tasks.
Key Features:
- Self-Supervised Pretraining: Employs a language-modeling self-supervised pretraining step on unlabeled protein sequences from Swiss-Prot to learn sequence representations.
- Task-Agnostic Representation: Produces generalizable, task-agnostic protein sequence embeddings that can be reused across different classification tasks.
- Fine-Tuning for Specific Tasks: Fine-tunes pretrained representations by updating model parameters to adapt to particular protein classification tasks.
Scientific Applications:
- Enzyme Class Prediction: Classifies proteins into enzyme classes using only amino acid sequence information.
- Gene Ontology Prediction: Predicts Gene Ontology terms covering biological processes, cellular components, and molecular functions from sequence data.
- Remote Homology and Fold Detection: Detects remote homologies and protein folds to support inference of evolutionary relationships and structural similarity.
Methodology:
Pretraining: self-supervised language-modeling on unlabeled Swiss-Prot sequences; Fine-tuning: supervised adaptation of the pretrained model to specific classification tasks by updating model parameters.
Topics
Details
- Programming Languages:
- Python, Shell
- Added:
- 1/18/2021
- Last Updated:
- 3/6/2021
Operations
Publications
Strodthoff N, Wagner P, Wenzel M, Samek W. UDSMProt: universal deep sequence models for protein classification. Bioinformatics. 2020;36(8):2401-2409. doi:10.1093/bioinformatics/btaa003. PMID:31913448. PMCID:PMC7178389.
PMID: 31913448
PMCID: PMC7178389
Funding: - Berlin Big Data Center: 01IS14013A
- Berlin Center for Machine Learning: 01IS18037I