AVADA

AVADA automates retrieval of variant-level pathogenic evidence from full-text primary literature to accelerate cataloging of disease-causing genetic variants for monogenic disease research and clinical diagnosis.


Key Features:

  • Machine learning and NLP: AVADA uses machine learning and natural language processing to scan and interpret full-text primary literature for variant evidence.
  • Variant-to-coordinate mapping: It extracts mentions of pathogenic genetic variants and converts them into precise genomic coordinates.
  • High retrieval efficiency: AVADA automatically retrieves approximately 60% of likely disease-causing variants listed in the Human Gene Mutation Database (HGMD), representing a 4.4-fold improvement over existing open-source automated variant extractors.
  • Comprehensive variant coverage: The system includes over 60,000 likely disease-causing variants present in HGMD but absent from ClinVar.
  • Diagnostic utility: In a cohort of 245 diagnosed patients, AVADA provided evidence for 18 additional diagnostic variants beyond those identified by ClinVar, compared with two additional variants found using the best current automated methods.
  • Full-text literature focus: AVADA operates on full-text primary literature and accelerates retrieval relative to manual searches through resources such as PubMed or Google Scholar.

Scientific Applications:

  • Pathogenic variant cataloging: Automated extraction of literature-based evidence to populate and expand variant databases such as HGMD and ClinVar.
  • Clinical diagnosis support: Provision of variant-level evidence to aid interpretation and diagnosis of monogenic disease patients.
  • Database curation: Identification of variants absent from ClinVar and provision of supporting literature citations for curation efforts.
  • Monogenic disease research: Facilitation of discovery and characterization of disease-causing variants and genes from the primary literature.

Methodology:

AVADA applies machine learning–based natural language processing to systematically scan and interpret full-text primary literature, extract mentions of pathogenic genetic variants, and map those mentions to genomic coordinates.

Topics

Details

Programming Languages:
Python
Added:
11/14/2019
Last Updated:
12/2/2020

Operations

Publications

Birgmeier J, Deisseroth CA, Hayward LE, Galhardo LM, Tierno AP, Jagadeesh KA, Stenson PD, Cooper DN, Bernstein JA, Haeussler M, Bejerano G. AVADA: toward automated pathogenic variant evidence retrieval directly from the full-text literature. Genetics in Medicine. 2020;22(2):362-370. doi:10.1038/s41436-019-0643-6. PMID:31467448. PMCID:PMC7301356.

PMID: 31467448
PMCID: PMC7301356
Funding: - EMBO: ALTF292-2011 - NIH/NHGRI: 5U41HG002371-15

Links