ProPheno
ProPheno aggregates mentions of human proteins and phenotypes from the complete corpora of Medline and PubMed Central Open Access to characterize protein-phenotype relationships for the study of rare and complex diseases.
Key Features:
- Extensive dataset: Constructed from the complete corpora of Medline and PubMed Central Open Access, aggregating mentions of human proteins and phenotypes.
- Textual co-occurrences: Captures co-occurrences of protein-phenotype pairs across textual spans, including sentences and paragraphs, to preserve contextual detail.
- Programmatic API: Exposes a RESTful API for programmatic queries and data retrieval.
Scientific Applications:
- Disease research: Supports analysis of protein-phenotype relationships relevant to rare and complex diseases, aiding identification of potential biomarkers and therapeutic targets.
- Biocuration: Provides a literature-derived source for validating and enriching databases with protein-phenotype associations.
- Literature mining: Enables locating and contextualizing articles that report protein-phenotype relationships to support hypothesis generation.
- Text-mining development: Serves as a dataset for training and evaluating predictive models that extract biological relationships from unstructured text.
Methodology:
Extracts protein-phenotype mentions from Medline and PubMed Central Open Access and analyzes co-occurrences within textual contexts (sentence and paragraph spans).
Topics
Details
- Added:
- 1/9/2020
- Last Updated:
- 12/6/2020
Operations
Publications
Pourreza Shahri M, Kahanda I. ProPheno 1.0: An online dataset for accelerating the complete characterization of the human protein-phenotype landscape in biomedical literature. Unknown Journal. 2019. doi:10.7287/peerj.preprints.27479v2.