cui2vec
cui2vec generates vector embeddings for 108,477 medical concepts by training unsupervised word-embedding models on multimodal clinical and biomedical corpora to represent relationships among medical concepts.
Key Features:
- R package: Implemented as an R package for generating and distributing embeddings.
- Multimodal corpus: Trained on 60 million insurance claims, 20 million clinical notes, and 1.7 million full-text biomedical journal articles.
- Vocabulary coverage: Produces embeddings for 108,477 medical concepts.
- Unsupervised word embeddings: Learns relationships among medical concepts using unsupervised word-embedding techniques.
- Co-occurrence fitting: Fits embeddings using co-occurrence data derived from the combined datasets.
- Data integration: Integrates disparate data types into a unified vector space representing medical concepts.
- Pre-trained embeddings: Provides pre-trained embeddings for downstream scientific analyses.
- Benchmark methodology: Includes a benchmark methodology specifically tailored to assess the statistical power of embeddings for medical concepts.
- Comparative performance: Demonstrates state-of-the-art performance in comparative analyses relative to prior methods.
Scientific Applications:
- Semantic characterization: Characterizing semantic relationships among medical concepts across claims, clinical notes, and biomedical literature.
- Clinical and biomedical NLP: Supporting natural language processing analyses of clinical notes and full-text biomedical journal articles.
- Statistical evaluation of embeddings: Assessing the statistical power of embedding-based analyses using the provided benchmark methodology.
- Method benchmarking: Benchmarking and comparing embedding approaches against previous methods.
Methodology:
Unsupervised learning of word embeddings fitted to co-occurrence data from 60 million insurance claims, 20 million clinical notes, and 1.7 million full-text biomedical journal articles, with multimodal data integration and validation via a benchmark assessing statistical power.
Topics
Details
- Tool Type:
- library
- Programming Languages:
- R
- Added:
- 1/14/2020
- Last Updated:
- 12/25/2020
Operations
Publications
Beam AL, et al. Clinical Concept Embeddings Learned from Massive Sources of Multimodal Medical Data. Pac Symp Biocomput. 2020; 25:295-306.
PMID: 31797605
PMCID: PMC6922053
Links
Issue tracker
https://github.com/beamandrew/cui2vec/issues