PubRunner

PubRunner automates periodic retrieval and processing of PubMed abstracts to maintain up-to-date biomedical text-mining corpora for analyses such as word2vec model updating and literature-based hypothesis generation.


Key Features:

  • Automated Scheduling: Automatically retrieves the most recent PubMed abstracts on a user-defined schedule.
  • Integration with Text Mining Tools: Incorporates user-specified text mining tools into the processing workflow for customized analyses.
  • Data Dissemination: Uploads processed results to public repositories such as FTP or Zenodo datasets.
  • Publicization of Results: Publishes the location of processed data on the PubRunner website.

Scientific Applications:

  • Up-to-date corpora for text mining: Maintains current literature corpora to support reproducible and contemporaneous text-mining analyses.
  • Improved understanding of biological interactions: Enables analysis of the latest findings to refine interpretation of complex biological interactions.
  • Faster hypothesis generation: Provides current literature-based evidence to accelerate generation of experimentally testable hypotheses.
  • word2vec re-training: Re-runs word2vec on the latest PubMed abstracts to generate current biomedical word vector representations.

Methodology:

Automated retrieval of recent PubMed abstracts on a user-defined schedule; processing of abstracts with a user-specified text mining tool; uploading processed data to public repositories such as FTP or Zenodo and publishing the data location on the PubRunner website.

Topics

Details

License:
MIT
Tool Type:
command-line tool, library
Programming Languages:
Python
Added:
8/13/2018
Last Updated:
12/10/2018

Operations

Publications

Anekalla KR, Courneya J, Fiorini N, Lever J, Muchow M, Busby B. PubRunner: A light-weight framework for updating text mining results. F1000Research. 2017;6:612. doi:10.12688/f1000research.11389.2. PMID:29152221. PMCID:PMC5664974.

Documentation

Links