MEME Suite
MEME Suite discovers and analyzes sequence motifs in DNA and protein sequences to identify conserved regulatory and functional elements.
Key Features:
- MEME (motif discovery): Uses expectation maximization to fit a two-component finite mixture model to unaligned sequences and discovers motifs by probabilistically erasing occurrences and repeating the search.
- GLAM2 (gapped motif discovery): Identifies motifs that contain insertions and deletions, enabling discovery of motifs with gaps.
- HMM modeling: Constructs Hidden Markov Models from motif models generated via the EM algorithm to model conserved regions within sequence families.
- Sequence scanning/search: Provides MAST, FIMO, and GLAM2SCAN to scan DNA and protein databases for motif occurrences, with MAST using MEME output and furnishing statistical measures for hits.
- Motif comparison: TOMTOM compares discovered motifs to motifs in existing motif databases.
- Functional association: GOMO associates motifs with Gene Ontology (GO) terms to infer potential functions.
- Output visualization: Produces sequence LOGOs for each discovered motif to visualize motif information content and positions.
Scientific Applications:
- Transcription Factor Binding Sites: Discovery of novel transcription factor binding sites in DNA sequences to study gene regulation.
- Protein Domains: Identification of conserved protein interaction domains and motifs relevant to protein function and interactions.
- Structural Analysis: Mapping motifs onto structural elements such as alpha-helices and beta-strands, exemplified by analyses of the short-chain alcohol dehydrogenase superfamily.
Methodology:
Expectation maximization fitting of a two-component finite mixture model to unaligned sequences with iterative probabilistic masking of motif occurrences; GLAM2 for gapped motif alignment; construction of HMMs from EM-derived motif models; database searches using MAST, FIMO, and GLAM2SCAN; motif comparison via TOMTOM and GO association via GOMO; generation of sequence LOGOs.
Topics
Details
- Tool Type:
- command-line tool, web application, workflow
- Operating Systems:
- Linux, Windows, Mac
- Added:
- 2/10/2017
- Last Updated:
- 2/4/2019
Operations
- Sequence motif discovery
- Nucleic acid feature detection
- Protein binding site prediction
- Sequence motif discovery
- Sequence motif discovery
- Nucleic acid feature detection
- Sequence motif discovery
- Enrichment analysis
- Statistical calculation
- Enrichment analysis
- Sequence profile alignment
- Visualisation
- Transcription factor binding site prediction
- Annotation
- Enrichment analysis
- Calculation
- Sequence motif recognition
- Sequence database search
- Sequence profile alignment
- Sequence clustering
- Sequence database search
- Sequence database search
- Sequence profile alignment
- Database search
- Sequence profile alignment
- Protein binding site prediction
- Sequence annotation
- Conversion
- Conversion
- Conversion
- Conversion
- Conversion
- Conversion
- Conversion
- Conversion
- Conversion
- Conversion
- Conversion
- Conversion
- Conversion
- Conversion
- Conversion
- Conversion
- Conversion
- Conversion
- Conversion
- Sequence conversion
- Sequence classification
- Statistical calculation
- Probabilistic data generation
- Statistical calculation
- Modelling and simulation
- Editing
- Loading
- Generation
- Statistical calculation
- Sequence generation
- Sequence composition calculation
- Formatting
- Annotation
- Formatting
- Statistical calculation
- Sequence motif discovery
- Sequence cutting
- Statistical calculation
- Comparison
- Phylogenetic tree analysis
- Sequence alignment editing
- Sequence alignment editing
- Phylogenetic footprinting / shadowing
Data Inputs & Outputs
Nucleic acid feature detection
Sequence motif discovery
Nucleic acid feature detection
Enrichment analysis
Sequence database search
Inputs
Outputs
Sequence database search
Inputs
Conversion
Outputs
Conversion
Outputs
Conversion
Conversion
Conversion
Outputs
Conversion
Annotation
Inputs
Outputs
Statistical calculation
Outputs
Sequence motif discovery
Sequence cutting
Statistical calculation
Inputs
Outputs
Comparison
Publications
Bailey TL, Boden M, Buske FA, Frith M, Grant CE, Clementi L, Ren J, Li WW, Noble WS. MEME SUITE: tools for motif discovery and searching. Nucleic Acids Research. 2009;37(Web Server):W202-W208. doi:10.1093/nar/gkp335. PMID:19458158. PMCID:PMC2703892.
Bailey TL and Elkan C. Fitting a mixture model by expectation maximization to discover motifs in biopolymers. Proc Int Conf Intell Syst Mol Biol. 1994; 2:28-36.
Bailey TL and Elkan C. The value of prior knowledge in discovering motifs with MEME. Proc Int Conf Intell Syst Mol Biol. 1995; 3:21-9.
Bailey TL and Gribskov M. The megaprior heuristic for discovering protein sequence patterns. Proc Int Conf Intell Syst Mol Biol. 1996; 4:15-24.
Grundy WN, et al. ParaMEME: a parallel implementation and a web interface for a DNA and protein motif discovery tool. Comput Appl Biosci. 1996; 12:303-10. doi: 10.1093/bioinformatics/12.4.303
Grundy WN, et al. Meta-MEME: motif-based hidden Markov models of protein families. Comput Appl Biosci. 1997; 13:397-406. doi: 10.1093/bioinformatics/13.4.397
Bailey TL, et al. An artificial intelligence approach to motif discovery in protein sequences: application to steriod dehydrogenases. J Steroid Biochem Mol Biol. 1997; 62:29-44. doi: 10.1016/s0960-0760(97)00013-7
Bailey TL, et al. MEME: discovering and analyzing DNA and protein sequence motifs. Nucleic Acids Res. 2006; 34:W369-73. doi: 10.1093/nar/gkl198