Enhancer-IF

Enhancer-IF identifies cell-specific enhancers in the human genome using an integrative machine learning framework.


Key Features:

  • Enhancer prediction: Identifies enhancer loci in human DNA sequences as regulatory DNA fragments that, when bound by transcription factors, enhance gene transcription.
  • Cell-type specificity: Incorporates cell-specific modelling to distinguish enhancers across different cellular contexts.
  • Addresses fixed-length sequence limitation: Explicitly targets limitations of using fixed-length sequences for model development.
  • Multi-cell datasets: Leverages datasets constructed from eight distinct cell types for training and evaluation.
  • Machine learning methods: Implements random forest, extremely randomized trees, multilayer perceptron, support vector machines, and extreme gradient boosting classifiers.
  • Sequence encodings: Trains models using seven different sequence encodings.
  • Baseline models: Generates 35 baseline models from combinations of five classifiers and seven encodings.
  • Meta-model construction: Integrates outputs of baseline models and feeds them back into the same classifiers to build meta-models.
  • Ensemble learning: Combines five meta-models into an ensemble to improve robustness.
  • Evaluation: Assesses performance on both training and independent datasets across various cell types.

Scientific Applications:

  • Genome-wide enhancer annotation: Predicts enhancer locations genome-wide in the human genome.
  • Cell-type specific enhancer discovery: Identifies enhancers specific to different cell types using models trained on eight cell-type datasets.
  • Functional annotation of enhancers: Supports annotation of potential enhancer functions and their physiological roles in cellular contexts.
  • Comparative model benchmarking: Enables comparison of predictive accuracy between baseline models, meta-models, and the final ensemble using training and independent datasets.

Methodology:

Five classifiers (random forest, extremely randomized trees, multilayer perceptron, support vector machines, extreme gradient boosting) are trained with seven sequence encodings to produce 35 baseline models; baseline outputs are integrated and re-input into the same classifiers to build five meta-models, which are combined into an ensemble and evaluated on training and independent datasets derived from eight distinct cell types.

Topics

Details

Tool Type:
workflow
Added:
11/27/2021
Last Updated:
11/27/2021

Operations

Publications

Basith S, Hasan MM, Lee G, Wei L, Manavalan B. Integrative machine learning framework for the identification of cell-specific enhancers from the human genome. Briefings in Bioinformatics. 2021;22(6). doi:10.1093/bib/bbab252. PMID:34226917.

PMID: 34226917
Funding: - National Natural Science Foundation of China: 2019R1I1A1A01062260, 2020R1A4A4079722, 2021R1A2C1014338, 62071278, 62072329