TFAnalysis
TFAnalysis predicts transcription factor binding sites and cell-type-specific interactions using ensemble random forest classifiers to support gene regulation analyses.
Key Features:
- Ensemble random forest models: Employs an ensemble learning strategy using random forest classifiers to predict TF binding sites across tissues and cell types.
- DNase1-seq data integration: Integrates DNase1-seq data and uses DNase1-seq peaks rather than raw signals to reduce false positives in binding predictions.
- Position Specific Energy Matrices (PSEMs): Uses PSEMs as transcription factor motif descriptors for model training and evaluation.
- ChIP-seq data integration: Uses transcription factor ChIP-seq data as gold-standard labels for training and evaluating predictive performance.
- Protein-protein interaction context: Feature-importance analyses reveal selection of motifs from TFs that are close interaction partners in protein-protein interaction networks.
Scientific Applications:
- Gene regulation studies: Predicts TF binding sites to aid elucidation of gene regulation mechanisms across various tissues and cell types.
- ENCODE-DREAM challenge evaluation: Developed and evaluated within the ENCODE-DREAM in vivo TF binding site prediction challenge.
Methodology:
Trains ensemble random forest classifiers using DNase1-seq peaks and PSEMs as input features, evaluates models against ChIP-seq data, and performs feature-importance analysis to identify cell-type-specific co-factors and motifs linked via protein-protein interaction networks.
Topics
Details
- License:
- MIT
- Programming Languages:
- R, Python
- Added:
- 11/14/2019
- Last Updated:
- 12/2/2020
Operations
Publications
Behjati Ardakani F, Schmidt F, Schulz MH. Predicting transcription factor binding using ensemble random forest models. F1000Research. 2019;7:1603. doi:10.12688/f1000research.16200.2.
Funding: - Cluster of Excellence on Multimodal Computing and Interaction: EXC248