PGxMine
PGxMine extracts pharmacogenomic associations from biomedical literature to identify relationships between genetic variants (DNA and protein changes, star alleles, and dbSNP identifiers) and chemicals for PharmGKB curation.
Key Features:
- Supervised machine learning pipeline: Employs a supervised machine learning approach to identify and extract associations between genetic variants (DNA and protein changes, star alleles, dbSNP identifiers) and chemicals from scientific texts.
- Large-scale literature mining: Mined over 7,170 papers resulting in 19,930 mentions of pharmacogenomic associations involving 452 chemicals and 2,426 variants.
- PharmGKB curation support: Identified novel associations with an evaluation showing 57 of the top 100 associations yielded 83 papers curatable by PharmGKB curators, plus 24 additional associations likely to lead to curatable content via citations.
- Integration with relation and workflow tools: Utilizes the Kindred relation classifier and integrates with the PubRunner project to process PubMed and accessible PubMed Central content.
Scientific Applications:
- Pharmacogenomic curation: Provides extracted association evidence to support curation efforts in PharmGKB.
- Precision medicine research: Supplies structured information on how genetic variants influence drug response and adverse effects to inform translational and clinical studies.
Methodology:
Applies text-mining techniques and supervised machine learning classifiers (including the Kindred relation classifier), implemented in Python3, and processes literature via PubRunner across PubMed and PubMed Central.
Topics
Details
- License:
- MIT
- Programming Languages:
- R, Python
- Added:
- 1/14/2020
- Last Updated:
- 1/9/2021
Operations
Publications
Lever J, et al. PGxMine: Text mining for curation of PharmGKB. Pac Symp Biocomput. 2020; 25:611-622.
PMID: 31797632
PMCID: PMC6917032
Links
Repository
https://github.com/jakelever/pgxmineRelated Tools
pharmgkb
Relation: includedIn