STREME
STREME identifies novel sequence motifs in DNA, RNA, and protein sequences to detect biologically relevant patterns such as binding sites.
Key Features:
- Algorithmic performance: Outperforms DREME, HOMER, MEME, Peak-motifs, ProSampler, and Weeder and accurately identifies motifs from in vivo DNA (ChIP-seq) and RNA (CLIP-seq) data validated against in vitro reference motifs.
- Large-scale capability: Handles datasets containing hundreds of thousands of sequences.
- Motif length range: Discovers short and long motifs ranging from 3 to 30 positions.
- Alphabet support: Operates on DNA, RNA, protein, and user-defined alphabets.
- Differential motif discovery: Supports comparative motif discovery between pairs of sequence datasets to identify condition-specific differences.
- Statistical significance: Provides an estimate of statistical significance for each discovered motif.
- Integration: Integrated with the MEME Suite for incorporation into broader sequence analysis workflows.
Scientific Applications:
- Binding site discovery: Identification of DNA- and RNA-binding protein motifs to study regulatory interactions.
- In vivo dataset analysis: Extraction of motifs from ChIP-seq and CLIP-seq experiments.
- Benchmarking and validation: Comparison and validation of discovered motifs against in vitro reference motifs.
- Comparative studies: Differential motif analysis to detect motif changes across conditions or sample groups.
Methodology:
Performs ab initio motif discovery using probabilistic and discrete models and computes statistical significance estimates for discovered motifs.
Topics
Details
- Tool Type:
- workflow
- Added:
- 12/6/2021
- Last Updated:
- 11/24/2024
Operations
Publications
Bailey TL. STREME: accurate and versatile sequence motif discovery. Bioinformatics. 2021;37(18):2834-2840. doi:10.1093/bioinformatics/btab203. PMID:33760053. PMCID:PMC8479671.