MariaOsmala
MariaOsmala predicts enhancers in the human genome by probabilistic modeling of chromatin feature patterns derived from ChIP-seq (next-generation sequencing) measurements of transcription factor binding sites and histone modifications.
Key Features:
- Input data: Uses ChIP-seq data from next-generation sequencing quantifying transcription factor binding sites and histone modifications.
- Probabilistic Modeling: Employs a statistical model to capture characteristic coverage patterns of chromatin features at enhancers and their variability.
- Probabilistic Distance Measures: Defines probabilistic distance measures that quantify similarity between genomic query regions and known enhancer coverage patterns.
- Kernel-Based Classifier Training: Trains a kernel-based classifier using probabilistic scores derived from both enhancer and non-enhancer samples.
- Robustness Across Cell Types: Demonstrates robust performance across different cell types, including those not used in training, with reports of outperforming existing state-of-the-art methods.
- Threshold Independence: Produces enhancer predictions that are less sensitive to the choice of prediction thresholds compared to other methods.
- Biological Validation and Novelty: Predictions have been computationally validated against transcriptional regulatory protein binding sites and include biologically relevant enhancers not identified by competing methods.
Scientific Applications:
- Functional genomics: Identification of enhancer elements to inform studies of gene regulation and regulatory architecture.
- Clinical genome interpretation: Expansion of enhancer catalogs to aid interpretation of noncoding variation in clinical and translational studies.
- Cross-cell-type regulatory analysis: Discovery and comparison of enhancers across diverse cell types, including cell types not present in training data.
Methodology:
Integrates ChIP-seq data with probabilistic/statistical modeling of chromatin coverage patterns; defines probabilistic distance measures between query regions and enhancer patterns; derives probabilistic scores for enhancer and non-enhancer samples and trains a kernel-based classifier; conducts computational validation against transcriptional regulatory protein binding sites.
Topics
Details
- License:
- MIT
- Tool Type:
- command-line tool
- Programming Languages:
- R, Shell
- Added:
- 1/9/2020
- Last Updated:
- 12/22/2020
Operations
Publications
Osmala M, Lähdesmäki H. Enhancer prediction in the human genome by probabilistic modelling of the chromatin feature patterns. Unknown Journal. 2019. doi:10.1101/804625.