BindSpace
BindSpace embeds DNA sequences and transcription factor labels into a shared representational space to enable multiclass prediction of TF–DNA binding.
Key Features:
- Joint sequence–label embedding: Trains embeddings that place DNA sequences and TF labels in the same representational space using binding data from hundreds of transcription factors and over one million DNA sequences.
- Multiclass binding prediction and validation: Produces multiclass predictions of TF binding with high performance that has been validated using both in vitro and in vivo data.
- Resolution of closely related TFs: Differentiates binding signals from closely related transcription factors to resolve similar binding motifs.
Scientific Applications:
- Transcription factor binding prediction: Predicts TF–DNA interactions for many TFs simultaneously through multiclass classification in the embedding space.
- Regulatory element identification: Aids identification of candidate regulatory elements by mapping sequence binding preferences to TF labels.
- Gene regulation studies: Supports analysis of gene regulation and discovery of regulatory mechanisms by linking sequences to TFs.
- Investigation of transcriptional dysregulation: Facilitates studies of genetic diseases linked to transcriptional dysregulation by predicting altered TF–DNA interactions.
Methodology:
Trains machine-learning embeddings on comprehensive TF binding datasets (hundreds of TFs, >1 million DNA sequences) to embed DNA sequences and TF labels in a shared space for multiclass binding prediction.
Topics
Details
- Added:
- 11/14/2019
- Last Updated:
- 12/5/2020
Operations
Publications
Yuan H, Kshirsagar M, Zamparo L, Lu Y, Leslie CS. BindSpace decodes transcription factor binding signals by large-scale sequence embedding. Nature Methods. 2019;16(9):858-861. doi:10.1038/s41592-019-0511-y. PMID:31406384. PMCID:PMC6717532.
PMID: 31406384
PMCID: PMC6717532
Funding: - U.S. Department of Health & Human Services | NIH | National Human Genome Research Institute: HG009395