SEME
SEME (Sampling with Expectation Maximization for Motif Elicitation) discovers de novo sequence motifs, particularly transcription factor (TF) binding sites, by integrating a probabilistic mixture model, importance sampling, and Expectation-Maximization to learn motif sequences, position preferences, and sequence-rank preferences.
Key Features:
- Probabilistic Mixture Model: Employs a pure probabilistic mixture model to represent motif and background binding characteristics.
- Expectation Maximization (EM) Algorithm: Uses EM to simultaneously learn sequence motifs, motif positions, and sequence-rank preferences from input sequences.
- Variable Motif Length Extension: Adapts motif length during analysis to detect motifs of varying sizes.
- Importance Sampling: Incorporates importance sampling to focus computation on high-probability regions of the sequence space.
- Position and Sequence-Rank Preferences: Models position preference and sequence-rank preference for motifs, including those of co-regulatory TFs (coTFs).
Scientific Applications:
- De novo TF binding-site discovery: Demonstrated improved motif discovery performance on 75 large-scale synthetic datasets, 32 metazoan compendium benchmark datasets, and 164 ChIP-Seq libraries.
- Co-regulated TF (coTF) motif identification: Identified a significantly higher number of correct coTF motifs across 15 ChIP-Seq libraries with predictions closely matching known motifs.
- Experimental validation: Learned position and sequence-rank preferences and motif predictions have been validated through additional ChIP-Seq experiments.
Methodology:
Applies a pure probabilistic mixture model with importance sampling and an Expectation-Maximization procedure that simultaneously learns sequence motifs, motif position preferences, and sequence-rank preferences, and incorporates a variable motif length extension for each coTF.
Topics
Details
- Tool Type:
- command-line tool
- Operating Systems:
- Linux, Windows, Mac
- Added:
- 8/3/2017
- Last Updated:
- 11/25/2024
Operations
Publications
Zhang Z, Chang CW, Hugo W, Cheung E, Sung W. Simultaneously Learning DNA Motif Along with Its Position and Sequence Rank Preferences Through Expectation Maximization Algorithm. Journal of Computational Biology. 2013;20(3):237-248. doi:10.1089/cmb.2012.0233. PMID:23461573.