ProPheno

ProPheno aggregates mentions of human proteins and phenotypes from the complete corpora of Medline and PubMed Central Open Access to characterize protein-phenotype relationships for the study of rare and complex diseases.


Key Features:

  • Extensive dataset: Constructed from the complete corpora of Medline and PubMed Central Open Access, aggregating mentions of human proteins and phenotypes.
  • Textual co-occurrences: Captures co-occurrences of protein-phenotype pairs across textual spans, including sentences and paragraphs, to preserve contextual detail.
  • Programmatic API: Exposes a RESTful API for programmatic queries and data retrieval.

Scientific Applications:

  • Disease research: Supports analysis of protein-phenotype relationships relevant to rare and complex diseases, aiding identification of potential biomarkers and therapeutic targets.
  • Biocuration: Provides a literature-derived source for validating and enriching databases with protein-phenotype associations.
  • Literature mining: Enables locating and contextualizing articles that report protein-phenotype relationships to support hypothesis generation.
  • Text-mining development: Serves as a dataset for training and evaluating predictive models that extract biological relationships from unstructured text.

Methodology:

Extracts protein-phenotype mentions from Medline and PubMed Central Open Access and analyzes co-occurrences within textual contexts (sentence and paragraph spans).

Topics

Details

Added:
1/9/2020
Last Updated:
12/6/2020

Operations

Publications

Pourreza Shahri M, Kahanda I. ProPheno 1.0: An online dataset for accelerating the complete characterization of the human protein-phenotype landscape in biomedical literature. Unknown Journal. 2019. doi:10.7287/peerj.preprints.27479v2.