HLA-SPREAD
HLA-SPREAD leverages Natural Language Processing (NLP) to curate and integrate associations between Human Leukocyte Antigens (HLA) and diseases, adverse drug reactions (ADRs), single nucleotide polymorphisms (SNPs), population annotations, and transplantation outcomes from PubMed abstracts.
Key Features:
- Automated information extraction: Processes approximately 24 million PubMed abstracts with NLP to extract HLA associations.
- Named Entity Recognition and semantic analysis: Applies Named Entity Recognition (NER) algorithms and semantic analysis to identify HLA alleles, diseases, ADRs, SNPs, and population entities.
- Hierarchical categorization: Organizes extracted information into 24 hierarchical groups for structured representation of HLA associations.
- Structured variant and population annotations: Extracts and structures data on SNPs, population annotations, and related resources associated with HLA findings.
- Population coverage: Aggregates associations reported across 112 countries and 32 ethnic groups.
- Clinical relevance mapping: Maps susceptible or risk alleles and clinically relevant biomarkers to autoimmune diseases, cancers, viral infections, skin disorders, transplantation outcomes, and hypersensitivity ADRs.
- False positive filtering and data structuring: Implements Python scripts to filter false positives and convert extracted records into a structured format.
Scientific Applications:
- Autoimmune Diseases: Identifying allelic and haplotypic HLA associations with autoimmune conditions.
- Cancer Research: Detecting HLA markers and allelic associations linked to cancer susceptibility.
- Viral Infections: Investigating associations between HLA alleles and viral disease outcomes.
- Transplantation Medicine: Assessing HLA alleles and markers that influence transplantation outcomes.
- Pharmacogenomics: Investigating hypersensitivity ADRs and other drug responses linked to specific HLA types.
Methodology:
Python scripts mined approximately 24 million PubMed abstracts; NLP extracted entities, NER algorithms and semantic analysis inferred HLA associations, false positives were filtered, and records were categorized into 24 hierarchical groups and structured formats.
Topics
Details
- Tool Type:
- web application
- Programming Languages:
- Python
- Added:
- 3/19/2021
- Last Updated:
- 3/31/2021
Operations
Publications
Dholakia D, Kalra A, Misir BR, Kanga U, Mukerji M. HLA-SPREAD: A Natural Language Processing based resource for curating HLA association from PubMed abstracts. Unknown Journal. 2021. doi:10.1101/2021.01.05.425409.