PubRunner
PubRunner automates periodic retrieval and processing of PubMed abstracts to maintain up-to-date biomedical text-mining corpora for analyses such as word2vec model updating and literature-based hypothesis generation.
Key Features:
- Automated Scheduling: Automatically retrieves the most recent PubMed abstracts on a user-defined schedule.
- Integration with Text Mining Tools: Incorporates user-specified text mining tools into the processing workflow for customized analyses.
- Data Dissemination: Uploads processed results to public repositories such as FTP or Zenodo datasets.
- Publicization of Results: Publishes the location of processed data on the PubRunner website.
Scientific Applications:
- Up-to-date corpora for text mining: Maintains current literature corpora to support reproducible and contemporaneous text-mining analyses.
- Improved understanding of biological interactions: Enables analysis of the latest findings to refine interpretation of complex biological interactions.
- Faster hypothesis generation: Provides current literature-based evidence to accelerate generation of experimentally testable hypotheses.
- word2vec re-training: Re-runs word2vec on the latest PubMed abstracts to generate current biomedical word vector representations.
Methodology:
Automated retrieval of recent PubMed abstracts on a user-defined schedule; processing of abstracts with a user-specified text mining tool; uploading processed data to public repositories such as FTP or Zenodo and publishing the data location on the PubRunner website.
Topics
Details
- License:
- MIT
- Tool Type:
- command-line tool, library
- Programming Languages:
- Python
- Added:
- 8/13/2018
- Last Updated:
- 12/10/2018
Operations
Publications
Anekalla KR, Courneya J, Fiorini N, Lever J, Muchow M, Busby B. PubRunner: A light-weight framework for updating text mining results. F1000Research. 2017;6:612. doi:10.12688/f1000research.11389.2. PMID:29152221. PMCID:PMC5664974.