PNIDB

PNIDB catalogs and annotates protein–nucleic acid interactions extracted from Protein Data Bank (PDB) structures and supplies machine-learning-based protein function predictions to support analysis of transcription, translation, DNA repair, and other protein–nucleic-acid-mediated cellular processes.


Key Features:

  • Data Source and Structure: Extracts structural data from the Protein Data Bank (PDB) and uses mmCIF keywords to classify nucleic acid-binding proteins and annotate entries.
  • Functional Classification: Categorizes proteins into 27 distinct classes, including transcription factors, immune system components, and structural proteins.
  • Predictive Analysis: Trains machine learning models on labeled sequences to predict protein functions, with reported 77.43% prediction accuracy validated by 10-fold cross-validation.
  • Entry Annotation: Provides per-entry annotations that summarize protein–nucleic acid interactions derived from PDB structures and mmCIF metadata.

Scientific Applications:

  • Transcriptional regulation: Supports research on transcriptional regulation by supplying structural interaction data and functional class labels for nucleic-acid-binding proteins.
  • Immune response studies: Enables investigation of immune system components and their nucleic acid interactions through classified and annotated entries.
  • Structural biology: Facilitates analysis of protein–nucleic acid complexes using experimentally solved PDB structures and derived annotations.
  • Viral infection and drug target research: Aids studies of viral infections and the identification of novel drug targets via function prediction of nucleic-acid-binding proteins.

Methodology:

Integrates structural data from PDB with functional annotations derived from mmCIF keywords; trains machine-learning algorithms on labeled sequence data; validates predictions by 10-fold cross-validation reporting 77.43% accuracy.

Topics

Details

Added:
1/18/2021
Last Updated:
1/24/2021

Operations

Publications

Xu L, Jiang S, Zou Q. An<i>in silico</i>approach to identification, categorization and prediction of nucleic acid binding proteins. Unknown Journal. 2020. doi:10.1101/2020.05.05.078741.