WordSpy

WordSpy identifies transcription factor binding motifs (TFBMs) within genomic promoter sequences to support discovery of cis-regulatory elements that control gene transcription.


Key Features:

  • Dictionary-based motif finding algorithm: Uses a predefined dictionary of motifs and a grammatical model combined with word counting to detect candidate TFBMs.
  • Word counting and statistical modeling: Quantifies occurrences of sequence words and applies a statistical model to analyze their frequency and distribution.
  • Genome-wide motif discovery: Scans large promoter datasets and can identify hundreds of motifs in a single run.
  • Integration of gene expression data: Incorporates gene expression information to improve discrimination between functional and spurious motifs.
  • Use of negative sequences: Employs negative sequence sets to identify discriminative motifs enriched in target promoters versus background.
  • Significance evaluation: Assesses motif significance by comparison to motifs found in randomly selected promoter sequences.
  • Comprehensive output: Produces an ordered list of putative motifs and associated regulatory sequences with annotated motif binding sites.

Scientific Applications:

  • Gene regulation studies: Identification of TFBMs to elucidate cis-regulatory elements involved in transcriptional control.
  • High-throughput promoter analysis: Genome-scale discovery of motifs across large promoter datasets.
  • Candidate prioritization for validation: Ranking of putative motifs and annotated binding sites to guide experimental follow-up.

Methodology:

The approach combines word counting, a statistical model using a motif dictionary and grammatical rules, integration of gene expression and negative sequence data, and significance assessment via comparison with randomly selected promoter sequences.

Topics

Details

Tool Type:
web application
Added:
3/24/2017
Last Updated:
11/25/2024

Operations

Publications

Wang G, Yu T, Zhang W. WordSpy: identifying transcription factor binding motifs by building a dictionary and learning a grammar. Nucleic Acids Research. 2005;33(Web Server):W412-W416. doi:10.1093/nar/gki492. PMID:15980501. PMCID:PMC1160252.