CCSeq
CCSeq identifies clusters of colocalized sequences along one-dimensional genomic structures to detect the spatial organization of transcription factor (TF) binding sites and infer potential TF complexes.
Key Features:
- Identification of Sequence Clusters: Detects clusters of sequences using genome annotation files and a user-specified cut-off distance to define colocalization along chromosomes.
- Analysis of Transcription Factor Complexes: Infers potential TF complexes and competitive occupancy of enhancer regions by analyzing spatial proximity of individual TF binding sequences.
- Application to Known TF Complexes: Has been applied to HSF1 homotrimer binding sequences and a NF-κB dimer composed of NFKB2 and RELB to validate spatial clustering hypotheses.
- Statistical Validation: Compares observed clusters to simulated random distributions, identifying, for example, 28 HSF1 trimer clusters on chromosome Y and 16 NFKB2+RELB dimer clusters on chromosome 17 with no clusters in five simulated random distributions for each TF set.
- Structural Pattern Analysis: Characterizes structural patterns within identified clusters to provide insights into genomic architecture associated with TF binding.
Scientific Applications:
- Transcription Factor Research: Elucidates spatial arrangements of TF binding sites to support studies of gene regulatory mechanisms.
- Genomic Annotation and Enhancer Analysis: Leverages genome annotation data to investigate enhancer regions and their potential TF occupancy patterns.
Methodology:
Analyzes genome annotation files using a cut-off distance parameter to define clusters, integrates position-weight matrix tools and genome annotation data from resources like PWMScan, compares observed clusters to simulated random distributions for statistical validation, and is implemented as an R package.
Topics
Details
- Programming Languages:
- R
- Added:
- 1/9/2020
- Last Updated:
- 12/10/2020
Operations
Publications
Golas S. CCSeq: Clusters of Colocalized Sequences. Unknown Journal. 2019. doi:10.1101/818385.
DOI: 10.1101/818385