MODER2
MODER2 discovers monomeric and dimeric transcription factor (TF)-binding motifs using first-order Markov models (adjacent dinucleotide matrices, ADMs) to model dependencies between adjacent nucleotides.
Key Features:
- First-order Markov models (ADMs): Models dependencies between adjacent nucleotides using adjacent dinucleotide matrices to capture dinucleotide context that position-specific probability matrices (PPMs) cannot.
- Monomeric and dimeric motif estimation: Simultaneously estimates models for monomeric and dimeric TF-binding motifs.
- ADM-based mixture model: Incorporates monomers and corresponding dimers within an ADM-based mixture framework.
- Dimeric structure modeling: Explicitly represents dimeric structure including spacing and orientation to capture cooperative dimerization effects.
- Expectation maximization learning: Fits model parameters using an expectation maximization algorithm from training data and seeds.
- Validation and benchmarking: Validated on HT-SELEX experiments and generated datasets and benchmarked against PPMs across 314 tested transcription factors or DNA-binding domains from bHLH, bZIP, ETS, and Homeodomain families.
- Implementation: Implemented in C++ with a dependency on the Boost library and tested on Linux platforms.
Scientific Applications:
- TF-binding motif discovery: Identification of monomeric and dimeric transcription factor-binding motifs from sequence data.
- Cooperative dimerization analysis: Investigation of cooperative binding effects by modeling spacing and orientation between monomers in dimers.
- High-throughput data analysis: Interpretation and modeling of HT-SELEX and generated motif datasets.
- Motif model benchmarking: Comparative evaluation of ADM mixture models versus PPMs across multiple TF families (bHLH, bZIP, ETS, Homeodomain).
Methodology:
Uses first-order Markov models (ADMs) within an ADM-based mixture model of monomers and dimers and fits parameters by expectation maximization from training data and seeds; validation performed on HT-SELEX and generated datasets and comparisons made to PPM models across 314 TFs from bHLH, bZIP, ETS, and Homeodomain families.
Topics
Details
- License:
- GPL-2.0
- Tool Type:
- command-line tool
- Programming Languages:
- C++, C, Python
- Added:
- 1/18/2021
- Last Updated:
- 2/26/2021
Operations
Publications
Toivonen J, Das PK, Taipale J, Ukkonen E. MODER2: first-order Markov modeling and discovery of monomeric and dimeric binding motifs. Bioinformatics. 2020;36(9):2690-2696. doi:10.1093/bioinformatics/btaa045. PMID:31999322. PMCID:PMC7203737.