MariaOsmala

MariaOsmala predicts enhancers in the human genome by probabilistic modeling of chromatin feature patterns derived from ChIP-seq (next-generation sequencing) measurements of transcription factor binding sites and histone modifications.


Key Features:

  • Input data: Uses ChIP-seq data from next-generation sequencing quantifying transcription factor binding sites and histone modifications.
  • Probabilistic Modeling: Employs a statistical model to capture characteristic coverage patterns of chromatin features at enhancers and their variability.
  • Probabilistic Distance Measures: Defines probabilistic distance measures that quantify similarity between genomic query regions and known enhancer coverage patterns.
  • Kernel-Based Classifier Training: Trains a kernel-based classifier using probabilistic scores derived from both enhancer and non-enhancer samples.
  • Robustness Across Cell Types: Demonstrates robust performance across different cell types, including those not used in training, with reports of outperforming existing state-of-the-art methods.
  • Threshold Independence: Produces enhancer predictions that are less sensitive to the choice of prediction thresholds compared to other methods.
  • Biological Validation and Novelty: Predictions have been computationally validated against transcriptional regulatory protein binding sites and include biologically relevant enhancers not identified by competing methods.

Scientific Applications:

  • Functional genomics: Identification of enhancer elements to inform studies of gene regulation and regulatory architecture.
  • Clinical genome interpretation: Expansion of enhancer catalogs to aid interpretation of noncoding variation in clinical and translational studies.
  • Cross-cell-type regulatory analysis: Discovery and comparison of enhancers across diverse cell types, including cell types not present in training data.

Methodology:

Integrates ChIP-seq data with probabilistic/statistical modeling of chromatin coverage patterns; defines probabilistic distance measures between query regions and enhancer patterns; derives probabilistic scores for enhancer and non-enhancer samples and trains a kernel-based classifier; conducts computational validation against transcriptional regulatory protein binding sites.

Topics

Details

License:
MIT
Tool Type:
command-line tool
Programming Languages:
R, Shell
Added:
1/9/2020
Last Updated:
12/22/2020

Operations

Publications

Osmala M, Lähdesmäki H. Enhancer prediction in the human genome by probabilistic modelling of the chromatin feature patterns. Unknown Journal. 2019. doi:10.1101/804625.