DeepFun

DeepFun predicts the functional impacts of non-coding genetic variants at single-nucleotide resolution using convolutional neural network-based deep learning and integrated epigenomic annotations.


Key Features:

  • Single-nucleotide resolution: Predicts variant effects at individual nucleotide positions in non-coding regions.
  • Non-coding focus: Targets non-coding regions of the human genome that account for over 90% of variants identified by GWAS.
  • Deep learning model: Employs a convolutional neural network (CNN) framework extensively evaluated for accuracy and reliability.
  • Epigenomic integration: Integrates annotations from ENCODE and Roadmap to inform predictions.
  • DNase I profiles: Incorporates 1548 DNase I accessibility profiles into the feature space.
  • Histone mark profiles: Incorporates 1536 histone mark profiles into the feature space.
  • Transcription factor profiles: Incorporates 4795 transcription factor binding profiles into the feature space.
  • Tissue and cell type specificity: Profiles span 225 distinct tissues or cell types to enable tissue- and cell type-specific assessments.
  • GWAS validation: Includes independent validations using datasets from various GWAS studies to assess predictive performance.
  • Motif visualization: Facilitates visualization of potential sequence motifs around variants.

Scientific Applications:

  • Variant prioritization in genetics: Prioritizes non-coding variants identified by GWAS for downstream analysis.
  • Functional genomics: Interprets regulatory effects of non-coding variants across tissues and cell types.
  • Disease mechanism studies: Associates non-coding variant impacts with disease-relevant tissues and pathways.
  • Personalized medicine and complex trait analysis: Provides detailed functional insights to inform studies of individual-level variant effects and complex genetic architectures.

Methodology:

Convolutional neural network-based deep learning trained on a feature space constructed from ENCODE and Roadmap epigenomic annotations (1548 DNase I profiles, 1536 histone mark profiles, 4795 transcription factor binding profiles across 225 tissues/cell types), with independent validation using GWAS datasets and motif visualization.

Topics

Details

Tool Type:
command-line tool, web application
Programming Languages:
R
Added:
9/8/2021
Last Updated:
11/24/2024

Operations

Publications

Pei G, Hu R, Jia P, Zhao Z. DeepFun: a deep learning sequence-based model to decipher non-coding variant effect in a tissue- and cell type-specific manner. Nucleic Acids Research. 2021;49(W1):W131-W139. doi:10.1093/nar/gkab429. PMID:34048560. PMCID:PMC8262726.

PMID: 34048560
PMCID: PMC8262726
Funding: - National Institutes of Health: R01DE030122, R01LM012806, R03DE027711 - Cancer Prevention and Research Institute of Texas: CPRIT RP180734, RP170668

Links