AVADA
AVADA automates retrieval of variant-level pathogenic evidence from full-text primary literature to accelerate cataloging of disease-causing genetic variants for monogenic disease research and clinical diagnosis.
Key Features:
- Machine learning and NLP: AVADA uses machine learning and natural language processing to scan and interpret full-text primary literature for variant evidence.
- Variant-to-coordinate mapping: It extracts mentions of pathogenic genetic variants and converts them into precise genomic coordinates.
- High retrieval efficiency: AVADA automatically retrieves approximately 60% of likely disease-causing variants listed in the Human Gene Mutation Database (HGMD), representing a 4.4-fold improvement over existing open-source automated variant extractors.
- Comprehensive variant coverage: The system includes over 60,000 likely disease-causing variants present in HGMD but absent from ClinVar.
- Diagnostic utility: In a cohort of 245 diagnosed patients, AVADA provided evidence for 18 additional diagnostic variants beyond those identified by ClinVar, compared with two additional variants found using the best current automated methods.
- Full-text literature focus: AVADA operates on full-text primary literature and accelerates retrieval relative to manual searches through resources such as PubMed or Google Scholar.
Scientific Applications:
- Pathogenic variant cataloging: Automated extraction of literature-based evidence to populate and expand variant databases such as HGMD and ClinVar.
- Clinical diagnosis support: Provision of variant-level evidence to aid interpretation and diagnosis of monogenic disease patients.
- Database curation: Identification of variants absent from ClinVar and provision of supporting literature citations for curation efforts.
- Monogenic disease research: Facilitation of discovery and characterization of disease-causing variants and genes from the primary literature.
Methodology:
AVADA applies machine learning–based natural language processing to systematically scan and interpret full-text primary literature, extract mentions of pathogenic genetic variants, and map those mentions to genomic coordinates.
Topics
Details
- Programming Languages:
- Python
- Added:
- 11/14/2019
- Last Updated:
- 12/2/2020
Operations
Publications
Birgmeier J, Deisseroth CA, Hayward LE, Galhardo LM, Tierno AP, Jagadeesh KA, Stenson PD, Cooper DN, Bernstein JA, Haeussler M, Bejerano G. AVADA: toward automated pathogenic variant evidence retrieval directly from the full-text literature. Genetics in Medicine. 2020;22(2):362-370. doi:10.1038/s41436-019-0643-6. PMID:31467448. PMCID:PMC7301356.