PepBCL

PepBCL predicts peptide-binding residues to identify protein-peptide interaction sites and support studies of protein function and drug discovery.


Key Features:

  • End-to-End Predictive Model: Operates as an end-to-end model that eliminates the need for handcrafted feature engineering and external preprocessing tools.
  • BERT-Based Framework: Uses a pre-trained protein language model (BERT-based) to automatically extract high-dimensional sequence representations.
  • Contrastive Learning Module: Incorporates a contrastive learning module to optimize feature representations and mitigate issues from imbalanced datasets.
  • Performance Superiority: Comparative benchmarking indicates that PepBCL significantly outperforms existing state-of-the-art methods in predictive performance.
  • Integration of Traditional and Learned Features: Combines traditional features with model-learned representations to enhance prediction accuracy.
  • Interpretable Analysis: Provides interpretable insights into binding residues, capturing both conserved and non-conserved sequential characteristics relevant to protein-peptide interactions.

Scientific Applications:

  • Protein function elucidation: Predicts binding sites on peptides to help elucidate mechanisms of protein function.
  • Drug discovery: Identifies potential peptide-binding sites on target proteins to support therapeutic development.

Methodology:

Uses a pre-trained protein language model (BERT-based) to learn sequence representations; applies a contrastive learning module to refine those representations and distinguish binding versus non-binding residues in imbalanced datasets; integrates traditional features with learned representations.

Topics

Details

License:
Not licensed
Cost:
Free of charge
Tool Type:
web application
Operating Systems:
Mac, Linux, Windows
Programming Languages:
Python
Added:
8/18/2022
Last Updated:
11/24/2024

Operations

Publications

Wang R, Jin J, Zou Q, Nakai K, Wei L. Predicting protein–peptide binding residues via interpretable deep learning. Bioinformatics. 2022;38(13):3351-3360. doi:10.1093/bioinformatics/btac352. PMID:35604077.

PMID: 35604077
Funding: - National Natural Science Foundation of China: 62071278

Links