MILLIPEDE
MILLIPEDE identifies transcription factor binding sites by integrating DNase digestion data with transcription factor binding specificity information to survey genomic locations of multiple TFs within a single experiment.
Key Features:
- Data integration: Integrates DNase digestion signals with transcription factor binding specificity information to evaluate genomic locations of multiple TFs.
- Performance superiority: Outperforms CENTIPEDE, marginally in human datasets and dramatically in yeast, increasing the average auROC across 20 TFs from 74% to 94%.
- Logistic regression framework: Uses a logistic regression statistical model as the core supervised learning component.
- Supervision versatility: Provides supervised, partially supervised, and completely unsupervised variants with performance close to supervised versions.
- Parameter efficiency: Requires at least an order of magnitude fewer parameters than CENTIPEDE.
Scientific Applications:
- Transcription factor binding site identification: Detects TF binding sites across multiple transcription factors using DNase digestion and specificity data.
- Gene regulation studies in human and yeast: Enables comparative analyses of transcriptional regulation mechanisms, with notably improved performance in yeast.
- High-throughput genomic analyses: Applicable to large-scale surveys of TF binding using DNase digestion experiments.
Methodology:
MILLIPEDE applies a logistic regression model that integrates DNase digestion signals with transcription factor binding specificity and is available in supervised, partially supervised, and unsupervised variants.
Topics
Details
- Tool Type:
- command-line tool
- Operating Systems:
- Linux, Windows, Mac
- Programming Languages:
- R
- Added:
- 8/3/2017
- Last Updated:
- 11/25/2024
Operations
Publications
Luo K and Hartemink AJ. Using DNase digestion data to accurately identify transcription factor binding sites. Pac Symp Biocomput. 2013; (unknown volume):80-91.
PMID: 23424114
PMCID: PMC3716004