discrover
discrover identifies discriminative sequence motifs in nucleic acid sequences using hidden Markov models and mutual information to distinguish motifs between positive and negative example sets.
Key Features:
- Discriminative learning with HMMs: Employs a discriminative learning approach based on hidden Markov models to identify sequence patterns that differ between datasets.
- Mutual information objective: Uses mutual information as the objective function to quantify the association between condition and motif occurrence.
- Performance and efficiency: Demonstrated higher accuracy and faster processing speeds in systematic comparisons with published motif-finding tools.
- Case studies across technologies: Applied to data types including ChIP-Seq, RIP-Chip, and PAR-CLIP, covering transcription factors in embryonic stem cells and RNA-binding proteins.
- Complex data configurations: Handles binary contrasts and more complex experimental configurations for genome- and transcriptome-scale analyses.
- Integration of repeat experiments: Makes use of available repeat experiments to enhance robustness and reliability of motif discovery.
Scientific Applications:
- Transcription factor motif discovery: Identification of transcription factor motifs from ChIP-Seq data, including studies in embryonic stem cells.
- RNA-binding protein motif discovery: Discovery of RNA-binding protein motifs from RIP-Chip and PAR-CLIP datasets.
- Alternative splicing factor analysis: Analysis of alternative splicing factors such as RBM10 to identify motifs relevant to splicing regulation.
- Genome- and transcriptome-scale motif analysis: Discriminative motif analysis across genome- and transcriptome-scale datasets and complex experimental designs.
- Robust motif identification with replicates: Use of repeat experiments to improve the robustness of discovered motifs.
Methodology:
Discriminative learning using hidden Markov models with mutual information as the objective function; systematic comparisons with published motif-finding tools; support for binary and multi-condition configurations and integration of repeat experiments.
Topics
Details
- License:
- GPL-3.0
- Cost:
- Free of charge
- Tool Type:
- command-line tool
- Operating Systems:
- Mac, Linux, Windows
- Programming Languages:
- C++, R
- Added:
- 3/16/2022
- Last Updated:
- 3/16/2022
Operations
Publications
Maaskola J, Rajewsky N. Binding site discovery from nucleic acid sequences by discriminative learning of hidden Markov models. Nucleic Acids Research. 2014;42(21):12995-13011. doi:10.1093/nar/gku1083. PMID:25389269. PMCID:PMC4245949.