EMCBOW-GPCR

EMCBOW-GPCR predicts G Protein-Coupled Receptors (GPCRs) from primary protein sequences using natural language processing-based word embeddings and machine learning to enable accurate GPCR identification.


Key Features:

  • Natural Language Processing (NLP) Integration: Applies NLP techniques to model protein sequence analysis analogously to text data processing.
  • Feature Extraction Models: Uses three distinct word-embedding models alongside a bag-of-words model to extract original features from GPCR primary sequences.
  • Deep Learning for Feature Enhancement: Employs deep learning algorithms to refine feature representations and reduce dimensionality.
  • Extreme Gradient Boosting (XGBoost): Classifies the processed features using Extreme Gradient Boosting for GPCR prediction.
  • Comparative Predictive Performance: Demonstrates performance that outperforms existing state-of-the-art methods on overall prediction metrics.

Scientific Applications:

  • GPCR Identification: Enables computational identification of G Protein-Coupled Receptors from sequence data as an alternative to experimental screening.
  • Drug Discovery and Development: Supports studies of receptor–ligand interactions and target identification in therapeutic development.
  • Membrane Protein Research: Facilitates analysis and screening of membrane protein sequence datasets.

Methodology:

GPCR primary sequences are represented using three word-embedding models and a bag-of-words model, processed by deep learning to extract, refine, and reduce feature dimensionality, and classified with Extreme Gradient Boosting (XGBoost).

Topics

Collections

Details

Cost:
Free of charge
Tool Type:
web application, workflow
Operating Systems:
Mac, Linux, Windows
Programming Languages:
Python
Added:
3/7/2022
Last Updated:
11/24/2024

Operations

Publications

Qiu W, Lv Z, Xiao X, Shao S, Lin H. EMCBOW-GPCR: A method for identifying G-protein coupled receptors based on word embedding and wordbooks. Computational and Structural Biotechnology Journal. 2021;19:4961-4969. doi:10.1016/j.csbj.2021.08.044. PMID:34527200. PMCID:PMC8437786.

PMID: 34527200
PMCID: PMC8437786
Funding: - Natural Science Foundation of Jiangxi Province: 20202BAB202007 - National Natural Science Foundation of China: 31760315, 31860312