WordSpy
WordSpy identifies transcription factor binding motifs (TFBMs) within genomic promoter sequences to support discovery of cis-regulatory elements that control gene transcription.
Key Features:
- Dictionary-based motif finding algorithm: Uses a predefined dictionary of motifs and a grammatical model combined with word counting to detect candidate TFBMs.
- Word counting and statistical modeling: Quantifies occurrences of sequence words and applies a statistical model to analyze their frequency and distribution.
- Genome-wide motif discovery: Scans large promoter datasets and can identify hundreds of motifs in a single run.
- Integration of gene expression data: Incorporates gene expression information to improve discrimination between functional and spurious motifs.
- Use of negative sequences: Employs negative sequence sets to identify discriminative motifs enriched in target promoters versus background.
- Significance evaluation: Assesses motif significance by comparison to motifs found in randomly selected promoter sequences.
- Comprehensive output: Produces an ordered list of putative motifs and associated regulatory sequences with annotated motif binding sites.
Scientific Applications:
- Gene regulation studies: Identification of TFBMs to elucidate cis-regulatory elements involved in transcriptional control.
- High-throughput promoter analysis: Genome-scale discovery of motifs across large promoter datasets.
- Candidate prioritization for validation: Ranking of putative motifs and annotated binding sites to guide experimental follow-up.
Methodology:
The approach combines word counting, a statistical model using a motif dictionary and grammatical rules, integration of gene expression and negative sequence data, and significance assessment via comparison with randomly selected promoter sequences.
Topics
Details
- Tool Type:
- web application
- Added:
- 3/24/2017
- Last Updated:
- 11/25/2024
Operations
Publications
Wang G, Yu T, Zhang W. WordSpy: identifying transcription factor binding motifs by building a dictionary and learning a grammar. Nucleic Acids Research. 2005;33(Web Server):W412-W416. doi:10.1093/nar/gki492. PMID:15980501. PMCID:PMC1160252.