MODER2

MODER2 discovers monomeric and dimeric transcription factor (TF)-binding motifs using first-order Markov models (adjacent dinucleotide matrices, ADMs) to model dependencies between adjacent nucleotides.


Key Features:

  • First-order Markov models (ADMs): Models dependencies between adjacent nucleotides using adjacent dinucleotide matrices to capture dinucleotide context that position-specific probability matrices (PPMs) cannot.
  • Monomeric and dimeric motif estimation: Simultaneously estimates models for monomeric and dimeric TF-binding motifs.
  • ADM-based mixture model: Incorporates monomers and corresponding dimers within an ADM-based mixture framework.
  • Dimeric structure modeling: Explicitly represents dimeric structure including spacing and orientation to capture cooperative dimerization effects.
  • Expectation maximization learning: Fits model parameters using an expectation maximization algorithm from training data and seeds.
  • Validation and benchmarking: Validated on HT-SELEX experiments and generated datasets and benchmarked against PPMs across 314 tested transcription factors or DNA-binding domains from bHLH, bZIP, ETS, and Homeodomain families.
  • Implementation: Implemented in C++ with a dependency on the Boost library and tested on Linux platforms.

Scientific Applications:

  • TF-binding motif discovery: Identification of monomeric and dimeric transcription factor-binding motifs from sequence data.
  • Cooperative dimerization analysis: Investigation of cooperative binding effects by modeling spacing and orientation between monomers in dimers.
  • High-throughput data analysis: Interpretation and modeling of HT-SELEX and generated motif datasets.
  • Motif model benchmarking: Comparative evaluation of ADM mixture models versus PPMs across multiple TF families (bHLH, bZIP, ETS, Homeodomain).

Methodology:

Uses first-order Markov models (ADMs) within an ADM-based mixture model of monomers and dimers and fits parameters by expectation maximization from training data and seeds; validation performed on HT-SELEX and generated datasets and comparisons made to PPM models across 314 TFs from bHLH, bZIP, ETS, and Homeodomain families.

Topics

Details

License:
GPL-2.0
Tool Type:
command-line tool
Programming Languages:
C++, C, Python
Added:
1/18/2021
Last Updated:
2/26/2021

Operations

Publications

Toivonen J, Das PK, Taipale J, Ukkonen E. MODER2: first-order Markov modeling and discovery of monomeric and dimeric binding motifs. Bioinformatics. 2020;36(9):2690-2696. doi:10.1093/bioinformatics/btaa045. PMID:31999322. PMCID:PMC7203737.

PMID: 31999322
PMCID: PMC7203737
Funding: - European Commission Framework Program 7 project SYSCOL: UE7-SYSCOL-258236 - Leverhulme Trust: VP1-2014-044 - Finnish CoE in Tumor Genetics Research: 312041