DISEASES
DISEASES aggregates and scores evidence for disease–gene associations by combining text-mined literature with manually curated data, cancer mutation datasets, and genome-wide association studies (GWAS).
Key Features:
- Dictionary-based named-entity recognition (NER): Uses a dictionary-based tagger to identify human genes and diseases in biomedical literature.
- Co-occurrence scoring: Applies a scoring scheme that evaluates co-occurrences within and between sentences to quantify associations.
- Evidence integration: Integrates text-mined associations with manually curated literature data, cancer mutation datasets, and GWAS results.
- Confidence scoring: Assigns confidence scores to each disease–gene association to enable comparison across evidence types.
- Performance metrics: Reported to recover approximately 50% of manually curated disease–gene associations with a false positive rate of 0.16%.
Scientific Applications:
- Hypothesis generation: Supports generation of hypotheses about genetic contributors to diseases by highlighting scored associations.
- Experimental validation: Aids validation and interpretation of experimental findings by providing integrated evidence for candidate genes.
- Target prioritization: Facilitates identification and prioritization of potential therapeutic targets based on combined evidence strength.
- Complex disease studies: Supports studies of complex diseases by aggregating diverse genetic evidence types.
Methodology:
Dictionary-based NER for human genes and diseases; co-occurrence scoring within and between sentences; integration with manually curated literature, cancer mutation datasets, and GWAS; assignment of confidence scores to associations.
Topics
Collections
Details
- License:
- CC-BY-4.0
- Maturity:
- Mature
- Cost:
- Free of charge
- Tool Type:
- web application
- Operating Systems:
- Linux, Windows, Mac
- Added:
- 2/23/2018
- Last Updated:
- 6/16/2020
Operations
Data Inputs & Outputs
Database search
Outputs
Database search
Publications
Pletscher-Frankild S, Pallejà A, Tsafou K, Binder JX, Jensen LJ. DISEASES: Text mining and data integration of disease–gene associations. Methods. 2015;74:83-89. doi:10.1016/j.ymeth.2014.11.020. PMID:25484339.
PMID: 25484339
Funding: - Novo Nordisk Foundation Center for Protein Research: NNF14CC0001
- European Union’s Seventh Framework Programme: FP7/2007-2013
Downloads
- Biological datahttps://diseases.jensenlab.org/DownloadsBulk download files in tab-delimited format.