NAGbinder

NAGbinder predicts N-acetylglucosamine (NAG) interacting residues in protein sequences to identify potential NAG-binding sites for studies of glycosylation, protein function, and disease mechanisms.


Key Features:

  • Data-Driven Approach: Derived from 231 nonredundant NAG-interacting protein chains extracted from the Protein Data Bank (PDB) with a maximum of 40% sequence identity, forming the basis for training, validation, and evaluation.
  • Model Development: Trained on a balanced dataset of 1,335 NAG-interacting and 1,335 noninteracting residues, with various window sizes tested, using Random Forest models built on binary residue profiles.
  • Performance Metrics: Achieved Matthews Correlation Coefficient (MCC) of 0.31 (training) and 0.25 (validation) with Area Under Receiver Operating Curve (AUROC) of 0.73 (training) and 0.70 (validation); on a realistic dataset of 1,335 interacting versus 47,198 noninteracting residues, reported MCCs of 0.26 (training) and 0.27 (validation) with AUROCs of 0.70 (training) and 0.71 (validation).
  • Practical Prediction Rate: When applied to a 1,000 amino acid sequence, approximately five out of ten predicted NAG-interacting residues are correctly identified.

Scientific Applications:

  • Glycosylation studies: Identifies putative NAG-binding sites to inform analyses of protein glycosylation.
  • Protein structure and function: Aids structural characterization and functional inference by highlighting likely NAG-interacting residues.
  • Disease and therapeutic research: Supports investigation of disease mechanisms involving NAG interactions and the development of therapeutics targeting glycoproteins.

Methodology:

Extraction of 231 nonredundant NAG-interacting chains from the PDB with ≤40% sequence identity; creation of a balanced dataset of 1,335 interacting and 1,335 noninteracting residues; testing of various window sizes; Random Forest model training using binary residue profiles; evaluation using MCC and AUROC including assessment on a realistic 1,335 versus 47,198 residue dataset.

Topics

Details

Tool Type:
web application
Programming Languages:
Python
Added:
1/9/2020
Last Updated:
11/24/2024

Operations

Publications

Patiyal S, Agrawal P, Kumar V, Dhall A, Kumar R, Mishra G, Raghava GP. NAGbinder: An approach for identifying N‐acetylglucosamine interacting residues of a protein from its primary sequence. Protein Science. 2019;29(1):201-210. doi:10.1002/pro.3761. PMID:31654438. PMCID:PMC6933864.

Links