STREME

STREME identifies novel sequence motifs in DNA, RNA, and protein sequences to detect biologically relevant patterns such as binding sites.


Key Features:

  • Algorithmic performance: Outperforms DREME, HOMER, MEME, Peak-motifs, ProSampler, and Weeder and accurately identifies motifs from in vivo DNA (ChIP-seq) and RNA (CLIP-seq) data validated against in vitro reference motifs.
  • Large-scale capability: Handles datasets containing hundreds of thousands of sequences.
  • Motif length range: Discovers short and long motifs ranging from 3 to 30 positions.
  • Alphabet support: Operates on DNA, RNA, protein, and user-defined alphabets.
  • Differential motif discovery: Supports comparative motif discovery between pairs of sequence datasets to identify condition-specific differences.
  • Statistical significance: Provides an estimate of statistical significance for each discovered motif.
  • Integration: Integrated with the MEME Suite for incorporation into broader sequence analysis workflows.

Scientific Applications:

  • Binding site discovery: Identification of DNA- and RNA-binding protein motifs to study regulatory interactions.
  • In vivo dataset analysis: Extraction of motifs from ChIP-seq and CLIP-seq experiments.
  • Benchmarking and validation: Comparison and validation of discovered motifs against in vitro reference motifs.
  • Comparative studies: Differential motif analysis to detect motif changes across conditions or sample groups.

Methodology:

Performs ab initio motif discovery using probabilistic and discrete models and computes statistical significance estimates for discovered motifs.

Topics

Details

Tool Type:
workflow
Added:
12/6/2021
Last Updated:
11/24/2024

Operations

Publications

Bailey TL. STREME: accurate and versatile sequence motif discovery. Bioinformatics. 2021;37(18):2834-2840. doi:10.1093/bioinformatics/btab203. PMID:33760053. PMCID:PMC8479671.

PMID: 33760053
PMCID: PMC8479671
Funding: - National Institutes of Health: R01 GM103544