EPIGENE

EPIGENE identifies active transcription units genome-wide by chromatin segmentation that correlates histone modifications with transcription activity, enabling TU detection independent of RNA-seq.


Key Features:

  • Chromatin segmentation: Correlates histone modifications with transcription activity to identify active transcription units (TUs) across the genome.
  • Multivariate Hidden Markov Model (HMM): Employs a constrained, semi-supervised multivariate HMM using a product of independent Bernoulli random variables to model combinations of histone modifications.
  • RNA-seq independence: Predicts TUs without relying on RNA-seq data, avoiding the requirement for large mRNA quantities.
  • Annotation concordance: Matches 93% of identified TUs to established gene annotations, with an additional 5% explained by microRNA annotations in HepG2 cells.
  • Comparative performance: Demonstrates higher TU prediction precision than RNA-seq-based approaches.
  • Novel TU discovery: Identifies novel TUs, including 381 genome-wide and 43 cell-specific units in tested cell lines such as K562, supported by RNA Polymerase II data.

Scientific Applications:

  • Functional and regulatory annotation: Identifies active TUs to elucidate functional and regulatory roles of genomic regions.
  • Unstable transcript detection: Facilitates study of unstable transcripts, including microRNA precursors, by detecting transcription independent of mRNA abundance.
  • Cell-line transcription mapping: Maps cell-specific transcription units across human cell lines (e.g., HepG2, K562).

Methodology:

Uses chromatin segmentation that correlates histone modifications with transcription activity and a constrained, semi-supervised multivariate HMM based on a product of independent Bernoulli random variables to analyze histone modification combinations; predictions are generated without RNA-seq and can be supported by RNA Polymerase II data.

Topics

Details

License:
MIT
Tool Type:
command-line tool
Programming Languages:
R
Added:
1/14/2020
Last Updated:
12/25/2020

Operations

Publications

Sahu A, Li N, Dunkel I, Chung H. EPIGENE: genome wide transcription unit annotation using a multivariate probabilistic model of histone modifications. Unknown Journal. 2019. doi:10.1101/2019.12.17.878454.

Links