BioVAE

BioVAE provides a large-scale pre-trained latent variable language model for biomedical text mining, integrating pre-trained language models with deep generative modeling to model and generate biomedical literature.


Key Features:

  • Domain-Specific Training: Trained on extensive volumes of biomedical literature using the OPTIMUS framework to capture biomedical text nuances.
  • Latent Variable Model: Implements a latent variable language model that uses deep generative models to model complex data distributions in biomedical text.
  • State-of-the-Art Performance: Demonstrates superior performance on multiple biomedical text mining benchmarks compared to other publicly available biomedical PLMs.
  • Sentence Generation: Generates biomedical sentences with higher accuracy than the original outputs from the OPTIMUS framework.

Scientific Applications:

  • Biomedical Literature Mining: Extracting relevant information and patterns from large collections of biomedical publications.
  • Data Annotation and Curation: Assisting annotation and curation of biomedical datasets by producing contextually accurate text.
  • Drug Discovery and Development: Supporting data analysis and text-generation tasks relevant to drug discovery and development research.

Methodology:

Trained via the OPTIMUS framework, integrating pre-trained language models (PLMs) with deep generative latent variable modeling on biomedical literature.

Topics

Details

License:
Apache-2.0
Cost:
Free of charge
Tool Type:
command-line tool
Programming Languages:
Shell
Added:
3/28/2022
Last Updated:
3/28/2022

Operations

Publications

Trieu H, Miwa M, Ananiadou S. BioVAE: a pre-trained latent variable language model for biomedical text mining. Bioinformatics. 2021;38(3):872-874. doi:10.1093/bioinformatics/btab702. PMID:34636886. PMCID:PMC8756089.

PMID: 34636886
PMCID: PMC8756089
Funding: - New Energy and Industrial Technology Development Organization: JPNP20006 - Alan Turing Institute and BBSRC: BB/P025684/1

Links