DRIMust

DRIMust identifies motifs in ranked sequence lists using suffix trees and a minimum hypergeometric statistical framework to detect over-represented variable gapped, long, and large-alphabet motifs and to report P-values and Position Specific Scoring Matrices (PSSMs).


Key Features:

  • Variable Gapped Motifs: Identifies motifs comprising two half-sites separated by a flexible-length gap.
  • Long Motifs and Large Alphabets: Searches for long motifs across large alphabets to accommodate complex biological sequences.
  • Efficient Enumeration with Suffix Trees: Utilizes suffix trees to efficiently enumerate motif candidates.
  • Minimum Hypergeometric Statistical Framework: Applies the minimum hypergeometric statistic to assess motif enrichment rigorously.
  • Data-Driven Thresholds: Dynamically determines thresholds for identifying over-represented motifs at the top of ranked lists.
  • Accurate P-value Assessment: Computes precise P-values for detected motifs.
  • Position Specific Scoring Matrix (PSSM): Outputs discovered motifs as individual motifs with associated P-values and as PSSMs.

Scientific Applications:

  • Sequence-based recognition (DNA, RNA, proteins): Supports study of sequence-based recognition processes involved in molecular regulation and signaling across DNA, RNA, and proteins.
  • High-throughput dataset analysis: Applied to analyze high-throughput measurement datasets for enriched sequence motifs.
  • Motif refinement and novel motif discovery: Used to refine known motifs such as the human estrogen receptor 1 motif and to identify variable-length motifs potentially associated with tyrosine phosphorylation.

Methodology:

Processes ranked lists of sequences in FASTA format; enumerates motif candidates using suffix trees; assesses motif significance with the minimum hypergeometric statistic; determines data-driven thresholds for enrichment; computes motif P-values and outputs motifs with Position Specific Scoring Matrices (PSSMs).

Topics

Details

Tool Type:
command-line tool
Operating Systems:
Linux, Windows
Programming Languages:
Java
Added:
12/18/2017
Last Updated:
11/25/2024

Operations

Publications

Leibovich L, Yakhini Z. Efficient motif search in ranked lists and applications to variable gap motifs. Nucleic Acids Research. 2012;40(13):5832-5847. doi:10.1093/nar/gks206. PMID:22416066. PMCID:PMC3401424.

Leibovich L, Paz I, Yakhini Z, Mandel-Gutfreund Y. DRIMust: a web server for discovering rank imbalanced motifs using suffix trees. Nucleic Acids Research. 2013;41(W1):W174-W179. doi:10.1093/nar/gkt407. PMID:23685432. PMCID:PMC3692051.

Documentation

Links