Enhancer-IF
Enhancer-IF identifies cell-specific enhancers in the human genome using an integrative machine learning framework.
Key Features:
- Enhancer prediction: Identifies enhancer loci in human DNA sequences as regulatory DNA fragments that, when bound by transcription factors, enhance gene transcription.
- Cell-type specificity: Incorporates cell-specific modelling to distinguish enhancers across different cellular contexts.
- Addresses fixed-length sequence limitation: Explicitly targets limitations of using fixed-length sequences for model development.
- Multi-cell datasets: Leverages datasets constructed from eight distinct cell types for training and evaluation.
- Machine learning methods: Implements random forest, extremely randomized trees, multilayer perceptron, support vector machines, and extreme gradient boosting classifiers.
- Sequence encodings: Trains models using seven different sequence encodings.
- Baseline models: Generates 35 baseline models from combinations of five classifiers and seven encodings.
- Meta-model construction: Integrates outputs of baseline models and feeds them back into the same classifiers to build meta-models.
- Ensemble learning: Combines five meta-models into an ensemble to improve robustness.
- Evaluation: Assesses performance on both training and independent datasets across various cell types.
Scientific Applications:
- Genome-wide enhancer annotation: Predicts enhancer locations genome-wide in the human genome.
- Cell-type specific enhancer discovery: Identifies enhancers specific to different cell types using models trained on eight cell-type datasets.
- Functional annotation of enhancers: Supports annotation of potential enhancer functions and their physiological roles in cellular contexts.
- Comparative model benchmarking: Enables comparison of predictive accuracy between baseline models, meta-models, and the final ensemble using training and independent datasets.
Methodology:
Five classifiers (random forest, extremely randomized trees, multilayer perceptron, support vector machines, extreme gradient boosting) are trained with seven sequence encodings to produce 35 baseline models; baseline outputs are integrated and re-input into the same classifiers to build five meta-models, which are combined into an ensemble and evaluated on training and independent datasets derived from eight distinct cell types.
Topics
Details
- Tool Type:
- workflow
- Added:
- 11/27/2021
- Last Updated:
- 11/27/2021
Operations
Publications
Basith S, Hasan MM, Lee G, Wei L, Manavalan B. Integrative machine learning framework for the identification of cell-specific enhancers from the human genome. Briefings in Bioinformatics. 2021;22(6). doi:10.1093/bib/bbab252. PMID:34226917.