MEME Suite

MEME Suite discovers and analyzes sequence motifs in DNA and protein sequences to identify conserved regulatory and functional elements.


Key Features:

  • MEME (motif discovery): Uses expectation maximization to fit a two-component finite mixture model to unaligned sequences and discovers motifs by probabilistically erasing occurrences and repeating the search.
  • GLAM2 (gapped motif discovery): Identifies motifs that contain insertions and deletions, enabling discovery of motifs with gaps.
  • HMM modeling: Constructs Hidden Markov Models from motif models generated via the EM algorithm to model conserved regions within sequence families.
  • Sequence scanning/search: Provides MAST, FIMO, and GLAM2SCAN to scan DNA and protein databases for motif occurrences, with MAST using MEME output and furnishing statistical measures for hits.
  • Motif comparison: TOMTOM compares discovered motifs to motifs in existing motif databases.
  • Functional association: GOMO associates motifs with Gene Ontology (GO) terms to infer potential functions.
  • Output visualization: Produces sequence LOGOs for each discovered motif to visualize motif information content and positions.

Scientific Applications:

  • Transcription Factor Binding Sites: Discovery of novel transcription factor binding sites in DNA sequences to study gene regulation.
  • Protein Domains: Identification of conserved protein interaction domains and motifs relevant to protein function and interactions.
  • Structural Analysis: Mapping motifs onto structural elements such as alpha-helices and beta-strands, exemplified by analyses of the short-chain alcohol dehydrogenase superfamily.

Methodology:

Expectation maximization fitting of a two-component finite mixture model to unaligned sequences with iterative probabilistic masking of motif occurrences; GLAM2 for gapped motif alignment; construction of HMMs from EM-derived motif models; database searches using MAST, FIMO, and GLAM2SCAN; motif comparison via TOMTOM and GO association via GOMO; generation of sequence LOGOs.

Topics

Details

Tool Type:
command-line tool, web application, workflow
Operating Systems:
Linux, Windows, Mac
Added:
2/10/2017
Last Updated:
2/4/2019

Operations

Data Inputs & Outputs

Sequence database search

Sequence database search

Conversion

Statistical calculation

Statistical calculation

Other operations do not define inputs or outputs.

Publications

Bailey TL, Boden M, Buske FA, Frith M, Grant CE, Clementi L, Ren J, Li WW, Noble WS. MEME SUITE: tools for motif discovery and searching. Nucleic Acids Research. 2009;37(Web Server):W202-W208. doi:10.1093/nar/gkp335. PMID:19458158. PMCID:PMC2703892.

Bailey TL and Elkan C. Fitting a mixture model by expectation maximization to discover motifs in biopolymers. Proc Int Conf Intell Syst Mol Biol. 1994; 2:28-36.

PMID: 7584402

Bailey TL and Elkan C. The value of prior knowledge in discovering motifs with MEME. Proc Int Conf Intell Syst Mol Biol. 1995; 3:21-9.

PMID: 7584439

Bailey TL and Gribskov M. The megaprior heuristic for discovering protein sequence patterns. Proc Int Conf Intell Syst Mol Biol. 1996; 4:15-24.

PMID: 8877500

Grundy WN, et al. ParaMEME: a parallel implementation and a web interface for a DNA and protein motif discovery tool. Comput Appl Biosci. 1996; 12:303-10. doi: 10.1093/bioinformatics/12.4.303

PMID: 8902357

Grundy WN, et al. Meta-MEME: motif-based hidden Markov models of protein families. Comput Appl Biosci. 1997; 13:397-406. doi: 10.1093/bioinformatics/13.4.397

PMID: 9283754

Bailey TL, et al. An artificial intelligence approach to motif discovery in protein sequences: application to steriod dehydrogenases. J Steroid Biochem Mol Biol. 1997; 62:29-44. doi: 10.1016/s0960-0760(97)00013-7

PMID: 9366496

Bailey TL, et al. MEME: discovering and analyzing DNA and protein sequence motifs. Nucleic Acids Res. 2006; 34:W369-73. doi: 10.1093/nar/gkl198

PMID: 16845028

Documentation