DISEASES

DISEASES aggregates and scores evidence for disease–gene associations by combining text-mined literature with manually curated data, cancer mutation datasets, and genome-wide association studies (GWAS).


Key Features:

  • Dictionary-based named-entity recognition (NER): Uses a dictionary-based tagger to identify human genes and diseases in biomedical literature.
  • Co-occurrence scoring: Applies a scoring scheme that evaluates co-occurrences within and between sentences to quantify associations.
  • Evidence integration: Integrates text-mined associations with manually curated literature data, cancer mutation datasets, and GWAS results.
  • Confidence scoring: Assigns confidence scores to each disease–gene association to enable comparison across evidence types.
  • Performance metrics: Reported to recover approximately 50% of manually curated disease–gene associations with a false positive rate of 0.16%.

Scientific Applications:

  • Hypothesis generation: Supports generation of hypotheses about genetic contributors to diseases by highlighting scored associations.
  • Experimental validation: Aids validation and interpretation of experimental findings by providing integrated evidence for candidate genes.
  • Target prioritization: Facilitates identification and prioritization of potential therapeutic targets based on combined evidence strength.
  • Complex disease studies: Supports studies of complex diseases by aggregating diverse genetic evidence types.

Methodology:

Dictionary-based NER for human genes and diseases; co-occurrence scoring within and between sentences; integration with manually curated literature, cancer mutation datasets, and GWAS; assignment of confidence scores to associations.

Topics

Collections

Details

License:
CC-BY-4.0
Maturity:
Mature
Cost:
Free of charge
Tool Type:
web application
Operating Systems:
Linux, Windows, Mac
Added:
2/23/2018
Last Updated:
6/16/2020

Operations

Data Inputs & Outputs

Publications

Pletscher-Frankild S, Pallejà A, Tsafou K, Binder JX, Jensen LJ. DISEASES: Text mining and data integration of disease–gene associations. Methods. 2015;74:83-89. doi:10.1016/j.ymeth.2014.11.020. PMID:25484339.

PMID: 25484339
Funding: - Novo Nordisk Foundation Center for Protein Research: NNF14CC0001 - European Union’s Seventh Framework Programme: FP7/2007-2013

Downloads