bioCADDIE
bioCADDIE provides indexed search and retrieval of publicly accessible biomedical datasets to enable dataset discovery and reuse.
Key Features:
- Information Retrieval System: Implements an advanced IR system tailored to retrieving relevant biomedical datasets from a growing repository based on user queries.
- Utilization of Unstructured Texts: Analyzes unstructured texts such as dataset titles and descriptions to improve retrieval accuracy.
- Medical Named Entity Extraction: Extracts and categorizes biomedical terms from dataset descriptions to enhance search precision.
- Deep Learning-Based Word Embeddings: Employs deep learning-derived word embeddings to expand queries with semantically related terms.
- Re-ranking Strategy: Applies a re-ranking strategy to prioritize datasets according to relevance to expanded queries.
Scientific Applications:
- Dataset discovery: Facilitates locating publicly accessible biomedical datasets relevant to specific research questions.
- Dataset reuse for research and validation: Supports reuse of discovered datasets for hypothesis testing and experimental validation in biomedical studies.
Methodology:
Analyzes unstructured text (titles, descriptions), applies medical named entity extraction, uses deep learning-based word embeddings for query expansion, employs re-ranking of search results, and was evaluated in the bioCADDIE Dataset Retrieval Challenge benchmarked against 11 baseline systems using inference Average Precision and inference normalized Discounted Cumulative Gain.
Topics
Details
- Programming Languages:
- Perl
- Added:
- 1/14/2020
- Last Updated:
- 12/9/2020
Operations
Publications
Wang Y, Rastegar-Mojarad M, Komandur-Elayavilli R, Liu H. Leveraging word embeddings and medical entity extraction for biomedical dataset retrieval using unstructured texts. Database. 2017;2017. doi:10.1093/database/bax091. PMID:31725862. PMCID:PMC7243926.