Pangaea

Pangaea extracts context-dependent gene–gene and gene–term relationships from PubMed biomedical literature for structured analysis of biological interactions.


Key Features:

  • Scalability and Extensibility: Scales to retrieval and processing of millions of PubMed articles by minimizing memory usage and optimizing parallel processing across multiple CPU cores, and supports incorporation of custom models.
  • Natural Language Processing (NLP): Applies advanced NLP methods to parse abstracts and extract structured relationships between genes and other biomedical terms.
  • Output Formats: Exports extracted relationships in CSV and JSON formats for integration with downstream analysis workflows.

Scientific Applications:

  • Gene Network Analysis: Constructs gene interaction networks from extracted gene–gene relationships to analyze biological pathways and processes.
  • Disease Mechanism Exploration: Identifies context-dependent relationships between genes and disease-related terms to aid elucidation of molecular disease mechanisms.
  • Drug Discovery and Development: Reveals gene–term associations that can inform drug target identification and repurposing hypotheses.

Methodology:

Automates downloading of PubMed articles, employs NLP methods to parse abstracts and extract structured relationships between genes and other terms, and uses memory-efficient data handling with parallel processing across multiple CPU cores.

Topics

Details

Added:
1/18/2021
Last Updated:
1/22/2021

Operations

Publications

Pirvan L, Samarajiwa SA. Pangaea: A modular and extensible collection of tools for mining context dependent gene relationships from the biomedical literature. Unknown Journal. 2020. doi:10.1101/2020.04.02.022517.

Documentation