cui2vec

cui2vec generates vector embeddings for 108,477 medical concepts by training unsupervised word-embedding models on multimodal clinical and biomedical corpora to represent relationships among medical concepts.


Key Features:

  • R package: Implemented as an R package for generating and distributing embeddings.
  • Multimodal corpus: Trained on 60 million insurance claims, 20 million clinical notes, and 1.7 million full-text biomedical journal articles.
  • Vocabulary coverage: Produces embeddings for 108,477 medical concepts.
  • Unsupervised word embeddings: Learns relationships among medical concepts using unsupervised word-embedding techniques.
  • Co-occurrence fitting: Fits embeddings using co-occurrence data derived from the combined datasets.
  • Data integration: Integrates disparate data types into a unified vector space representing medical concepts.
  • Pre-trained embeddings: Provides pre-trained embeddings for downstream scientific analyses.
  • Benchmark methodology: Includes a benchmark methodology specifically tailored to assess the statistical power of embeddings for medical concepts.
  • Comparative performance: Demonstrates state-of-the-art performance in comparative analyses relative to prior methods.

Scientific Applications:

  • Semantic characterization: Characterizing semantic relationships among medical concepts across claims, clinical notes, and biomedical literature.
  • Clinical and biomedical NLP: Supporting natural language processing analyses of clinical notes and full-text biomedical journal articles.
  • Statistical evaluation of embeddings: Assessing the statistical power of embedding-based analyses using the provided benchmark methodology.
  • Method benchmarking: Benchmarking and comparing embedding approaches against previous methods.

Methodology:

Unsupervised learning of word embeddings fitted to co-occurrence data from 60 million insurance claims, 20 million clinical notes, and 1.7 million full-text biomedical journal articles, with multimodal data integration and validation via a benchmark assessing statistical power.

Topics

Details

Tool Type:
library
Programming Languages:
R
Added:
1/14/2020
Last Updated:
12/25/2020

Operations

Publications

Beam AL, et al. Clinical Concept Embeddings Learned from Massive Sources of Multimodal Medical Data. Pac Symp Biocomput. 2020; 25:295-306.

PMID: 31797605
PMCID: PMC6922053

Links