TFAnalysis

TFAnalysis predicts transcription factor binding sites and cell-type-specific interactions using ensemble random forest classifiers to support gene regulation analyses.


Key Features:

  • Ensemble random forest models: Employs an ensemble learning strategy using random forest classifiers to predict TF binding sites across tissues and cell types.
  • DNase1-seq data integration: Integrates DNase1-seq data and uses DNase1-seq peaks rather than raw signals to reduce false positives in binding predictions.
  • Position Specific Energy Matrices (PSEMs): Uses PSEMs as transcription factor motif descriptors for model training and evaluation.
  • ChIP-seq data integration: Uses transcription factor ChIP-seq data as gold-standard labels for training and evaluating predictive performance.
  • Protein-protein interaction context: Feature-importance analyses reveal selection of motifs from TFs that are close interaction partners in protein-protein interaction networks.

Scientific Applications:

  • Gene regulation studies: Predicts TF binding sites to aid elucidation of gene regulation mechanisms across various tissues and cell types.
  • ENCODE-DREAM challenge evaluation: Developed and evaluated within the ENCODE-DREAM in vivo TF binding site prediction challenge.

Methodology:

Trains ensemble random forest classifiers using DNase1-seq peaks and PSEMs as input features, evaluates models against ChIP-seq data, and performs feature-importance analysis to identify cell-type-specific co-factors and motifs linked via protein-protein interaction networks.

Topics

Details

License:
MIT
Programming Languages:
R, Python
Added:
11/14/2019
Last Updated:
12/2/2020

Operations

Publications

Behjati Ardakani F, Schmidt F, Schulz MH. Predicting transcription factor binding using ensemble random forest models. F1000Research. 2019;7:1603. doi:10.12688/f1000research.16200.2.

Funding: - Cluster of Excellence on Multimodal Computing and Interaction: EXC248