LOGO
LOGO interprets non-coding regions of the human genome using a self-attention based contextualized pre-trained language model that integrates convolutional modules and self-supervised learning to enable base-resolution sequence labeling and variant prioritization.
Key Features:
- Lightweight Architecture: The model uses two self-attention layers totaling 1 million parameters for a compact representation.
- Self-Supervised Learning: Learns bidirectional representations from the unlabelled human reference genome via self-supervision.
- Contextualized Sequence Representation: Captures contextual information within genomic sequences to support sequence labeling and variant prioritization.
- Variant Prioritization: Implements a specialized input encoding for alternative alleles combined with a convolutional module to prioritize non-coding variants associated with human disease.
- Performance Metrics: Reports a 15% absolute improvement in promoter identification, up to a 4.5% absolute enhancement in enhancer-promoter interaction prediction, and state-of-the-art predictive power on thousands of chromatin features while using ~3% of the parameters of DeepSEA and ~1% of recent BERT-based DNA language models.
Scientific Applications:
- Base-resolution Sequence Labeling: Enables global sequence labeling at base resolution across the genome.
- Enhancer-Promoter Interaction Prediction: Improves prediction of enhancer-promoter interactions, with up to 4.5% absolute performance gains reported.
- Chromatin Feature Prediction: Provides state-of-the-art predictive power on thousands of chromatin features.
- Disease GWAS Interpretation: Applied to interpret type 2 diabetes (T2D) GWAS signals to infer underlying regulatory mechanisms.
Methodology:
Uses a self-attention based contextualized pre-trained language model with two self-attention layers (~1 million parameters), integrates convolutional modules, applies self-supervised bidirectional representation learning on the unlabelled human reference genome, and employs a specialized input encoding for alternative alleles.
Topics
Details
- License:
- Not licensed
- Cost:
- Free of charge
- Tool Type:
- command-line tool
- Operating Systems:
- Mac, Linux, Windows
- Programming Languages:
- Python
- Added:
- 8/28/2022
- Last Updated:
- 11/24/2024
Operations
Publications
Yang M, Huang L, Huang H, Tang H, Zhang N, Yang H, Wu J, Mu F. Integrating convolution and self-attention improves language model of human genome for interpreting non-coding regions at base-resolution. Nucleic Acids Research. 2022;50(14):e81-e81. doi:10.1093/nar/gkac326. PMID:35536244. PMCID:PMC9371931.