RLIMS-P
RLIMS-P 2.0 extracts information on protein phosphorylation from biomedical literature by identifying protein kinases, substrates, and specific phosphorylation sites to support curation and analysis of post-translational modifications.
Key Features:
- Rule-based text mining: Implements a rule-based text-mining program to identify phosphorylation-related events in text.
- Entity targets: Detects protein kinases, substrates, and phosphorylation sites as explicit biological entities.
- NLP and IE modules: Employs natural language processing (NLP) and information extraction (IE) modules to process literature.
- Abstract and full-text processing: Processes both abstracts and full-text articles for comprehensive extraction.
- Expanded expression coverage: Handles a variety of textual expressions found in scientific literature to improve accuracy and generalizability.
- Performance on abstracts: Achieved F-scores of 0.91 (kinases), 0.92 (substrates), and 0.95 (phosphorylation sites) on annotated abstract corpora.
- Performance on full-text: Achieved F-scores of 0.88 (kinases), 0.91 (substrates), and 0.92 (phosphorylation sites) on annotated full-text corpora.
- BioNLP-ST GE 2013 evaluation: Achieved an F-score of 0.87 for phosphorylation core tasks on the 2013 BioNLP-ST GE corpus.
- Scalability: Validated by full-scale processing of all abstracts in MEDLINE and all articles in the PubMed Central Open Access Subset.
- Adaptability to other PTMs: Designed to be extendable to other major post-translational modification (PTM) types.
Scientific Applications:
- Biocuration: Enables extraction of phosphorylation events for structured database curation of PTMs.
- Proteomics research: Supports analysis of kinase–substrate relationships and site-specific phosphorylation in proteomics studies.
- Knowledge discovery: Facilitates literature-based discovery of phosphorylation-related mechanisms and hypotheses.
- Benchmarking and method development: Provides evaluated outputs for comparison in text-mining and information-extraction research.
Methodology:
Uses a rule-based text-mining approach with NLP and IE modules to process abstracts and full-text articles, evaluated on annotated corpora including the BioNLP-ST GE 2013 dataset and full-scale MEDLINE and PubMed Central Open Access Subset processing.
Topics
Details
- Tool Type:
- command-line tool
- Operating Systems:
- Linux, Windows, Mac
- Added:
- 12/18/2017
- Last Updated:
- 11/25/2024
Operations
Publications
Torii M, Arighi CN, Li G, Wang Q, Wu CH, Vijay-Shanker K. RLIMS-P 2.0: A Generalizable Rule-Based Information Extraction System for Literature Mining of Protein Phosphorylation Information. IEEE/ACM Transactions on Computational Biology and Bioinformatics. 2015;12(1):17-29. doi:10.1109/tcbb.2014.2372765. PMID:26357075. PMCID:PMC4568560.