RLIMS-P

RLIMS-P 2.0 extracts information on protein phosphorylation from biomedical literature by identifying protein kinases, substrates, and specific phosphorylation sites to support curation and analysis of post-translational modifications.


Key Features:

  • Rule-based text mining: Implements a rule-based text-mining program to identify phosphorylation-related events in text.
  • Entity targets: Detects protein kinases, substrates, and phosphorylation sites as explicit biological entities.
  • NLP and IE modules: Employs natural language processing (NLP) and information extraction (IE) modules to process literature.
  • Abstract and full-text processing: Processes both abstracts and full-text articles for comprehensive extraction.
  • Expanded expression coverage: Handles a variety of textual expressions found in scientific literature to improve accuracy and generalizability.
  • Performance on abstracts: Achieved F-scores of 0.91 (kinases), 0.92 (substrates), and 0.95 (phosphorylation sites) on annotated abstract corpora.
  • Performance on full-text: Achieved F-scores of 0.88 (kinases), 0.91 (substrates), and 0.92 (phosphorylation sites) on annotated full-text corpora.
  • BioNLP-ST GE 2013 evaluation: Achieved an F-score of 0.87 for phosphorylation core tasks on the 2013 BioNLP-ST GE corpus.
  • Scalability: Validated by full-scale processing of all abstracts in MEDLINE and all articles in the PubMed Central Open Access Subset.
  • Adaptability to other PTMs: Designed to be extendable to other major post-translational modification (PTM) types.

Scientific Applications:

  • Biocuration: Enables extraction of phosphorylation events for structured database curation of PTMs.
  • Proteomics research: Supports analysis of kinase–substrate relationships and site-specific phosphorylation in proteomics studies.
  • Knowledge discovery: Facilitates literature-based discovery of phosphorylation-related mechanisms and hypotheses.
  • Benchmarking and method development: Provides evaluated outputs for comparison in text-mining and information-extraction research.

Methodology:

Uses a rule-based text-mining approach with NLP and IE modules to process abstracts and full-text articles, evaluated on annotated corpora including the BioNLP-ST GE 2013 dataset and full-scale MEDLINE and PubMed Central Open Access Subset processing.

Topics

Details

Tool Type:
command-line tool
Operating Systems:
Linux, Windows, Mac
Added:
12/18/2017
Last Updated:
11/25/2024

Operations

Publications

Torii M, Arighi CN, Li G, Wang Q, Wu CH, Vijay-Shanker K. RLIMS-P 2.0: A Generalizable Rule-Based Information Extraction System for Literature Mining of Protein Phosphorylation Information. IEEE/ACM Transactions on Computational Biology and Bioinformatics. 2015;12(1):17-29. doi:10.1109/tcbb.2014.2372765. PMID:26357075. PMCID:PMC4568560.

PMID: 26357075
PMCID: PMC4568560
Funding: - National Institutes of Health: G08LM010720 - National Science Foundation: ABI-1062520

Documentation

Links