Pangaea
Pangaea extracts context-dependent gene–gene and gene–term relationships from PubMed biomedical literature for structured analysis of biological interactions.
Key Features:
- Scalability and Extensibility: Scales to retrieval and processing of millions of PubMed articles by minimizing memory usage and optimizing parallel processing across multiple CPU cores, and supports incorporation of custom models.
- Natural Language Processing (NLP): Applies advanced NLP methods to parse abstracts and extract structured relationships between genes and other biomedical terms.
- Output Formats: Exports extracted relationships in CSV and JSON formats for integration with downstream analysis workflows.
Scientific Applications:
- Gene Network Analysis: Constructs gene interaction networks from extracted gene–gene relationships to analyze biological pathways and processes.
- Disease Mechanism Exploration: Identifies context-dependent relationships between genes and disease-related terms to aid elucidation of molecular disease mechanisms.
- Drug Discovery and Development: Reveals gene–term associations that can inform drug target identification and repurposing hypotheses.
Methodology:
Automates downloading of PubMed articles, employs NLP methods to parse abstracts and extract structured relationships between genes and other terms, and uses memory-efficient data handling with parallel processing across multiple CPU cores.
Topics
Details
- Added:
- 1/18/2021
- Last Updated:
- 1/22/2021
Operations
Publications
Pirvan L, Samarajiwa SA. Pangaea: A modular and extensible collection of tools for mining context dependent gene relationships from the biomedical literature. Unknown Journal. 2020. doi:10.1101/2020.04.02.022517.