LOGO

LOGO interprets non-coding regions of the human genome using a self-attention based contextualized pre-trained language model that integrates convolutional modules and self-supervised learning to enable base-resolution sequence labeling and variant prioritization.


Key Features:

  • Lightweight Architecture: The model uses two self-attention layers totaling 1 million parameters for a compact representation.
  • Self-Supervised Learning: Learns bidirectional representations from the unlabelled human reference genome via self-supervision.
  • Contextualized Sequence Representation: Captures contextual information within genomic sequences to support sequence labeling and variant prioritization.
  • Variant Prioritization: Implements a specialized input encoding for alternative alleles combined with a convolutional module to prioritize non-coding variants associated with human disease.
  • Performance Metrics: Reports a 15% absolute improvement in promoter identification, up to a 4.5% absolute enhancement in enhancer-promoter interaction prediction, and state-of-the-art predictive power on thousands of chromatin features while using ~3% of the parameters of DeepSEA and ~1% of recent BERT-based DNA language models.

Scientific Applications:

  • Base-resolution Sequence Labeling: Enables global sequence labeling at base resolution across the genome.
  • Enhancer-Promoter Interaction Prediction: Improves prediction of enhancer-promoter interactions, with up to 4.5% absolute performance gains reported.
  • Chromatin Feature Prediction: Provides state-of-the-art predictive power on thousands of chromatin features.
  • Disease GWAS Interpretation: Applied to interpret type 2 diabetes (T2D) GWAS signals to infer underlying regulatory mechanisms.

Methodology:

Uses a self-attention based contextualized pre-trained language model with two self-attention layers (~1 million parameters), integrates convolutional modules, applies self-supervised bidirectional representation learning on the unlabelled human reference genome, and employs a specialized input encoding for alternative alleles.

Topics

Details

License:
Not licensed
Cost:
Free of charge
Tool Type:
command-line tool
Operating Systems:
Mac, Linux, Windows
Programming Languages:
Python
Added:
8/28/2022
Last Updated:
11/24/2024

Operations

Publications

Yang M, Huang L, Huang H, Tang H, Zhang N, Yang H, Wu J, Mu F. Integrating convolution and self-attention improves language model of human genome for interpreting non-coding regions at base-resolution. Nucleic Acids Research. 2022;50(14):e81-e81. doi:10.1093/nar/gkac326. PMID:35536244. PMCID:PMC9371931.

PMID: 35536244
PMCID: PMC9371931
Funding: - Guangdong Provincial Academician Workstation of BGI Synthetic Genomics: 2017B090904014 - Program of Shanghai Academic Research Leader: 20XD1401100 - Program for Outstanding Medical Academic Leader: 2019LJ01