BioSeq-BLM
BioSeq-BLM analyzes DNA, RNA, and protein sequences using 155 biological language models to extract linguistic properties for sequence representation and predictive modeling.
Key Features:
- Biological Language Models (BLMs): Integrates 155 BLMs tailored for DNA, RNA, and protein sequences to identify and analyze linguistic patterns in biological data.
- Automated Sequence Representation: Extends BLMs into an automated system for representing and analyzing sequence data.
- Predictor Development and Performance: Produces predictors whose experimental performance is reported as comparable to or surpassing existing state-of-the-art methods in biological sequence prediction.
Scientific Applications:
- Genomic and Transcriptomic Analysis: Applies BLM-derived representations to investigate genomic structures and transcriptomic dynamics.
- Protein Function Prediction: Uses sequence-based linguistic features to inform predictions of protein function.
- Comparative Genomics: Enables comparative analyses across species by comparing sequence linguistic patterns.
Methodology:
Uses 155 biological language models to extract and interpret linguistic properties of DNA, RNA, and protein sequences and to generate automated sequence representations and predictive models.
Topics
Details
- Cost:
- Free of charge
- Tool Type:
- web application
- Operating Systems:
- Mac, Linux, Windows
- Added:
- 2/6/2022
- Last Updated:
- 11/24/2024
Operations
Publications
Li H, Pang Y, Liu B. BioSeq-BLM: a platform for analyzing DNA, RNA and protein sequences based on biological language models. Nucleic Acids Research. 2021;49(22):e129-e129. doi:10.1093/nar/gkab829. PMID:34581805. PMCID:PMC8682797.
DOI: 10.1093/nar/gkab829
PMID: 34581805
PMCID: PMC8682797
Funding: - National Key Research and Development Program of China: 2018AAA0100100
- National Natural Science Foundation of China: 61732012, 61822306, 61861146002
- Beijing Natural Science Foundation: JQ19019