CCSeq

CCSeq identifies clusters of colocalized sequences along one-dimensional genomic structures to detect the spatial organization of transcription factor (TF) binding sites and infer potential TF complexes.


Key Features:

  • Identification of Sequence Clusters: Detects clusters of sequences using genome annotation files and a user-specified cut-off distance to define colocalization along chromosomes.
  • Analysis of Transcription Factor Complexes: Infers potential TF complexes and competitive occupancy of enhancer regions by analyzing spatial proximity of individual TF binding sequences.
  • Application to Known TF Complexes: Has been applied to HSF1 homotrimer binding sequences and a NF-κB dimer composed of NFKB2 and RELB to validate spatial clustering hypotheses.
  • Statistical Validation: Compares observed clusters to simulated random distributions, identifying, for example, 28 HSF1 trimer clusters on chromosome Y and 16 NFKB2+RELB dimer clusters on chromosome 17 with no clusters in five simulated random distributions for each TF set.
  • Structural Pattern Analysis: Characterizes structural patterns within identified clusters to provide insights into genomic architecture associated with TF binding.

Scientific Applications:

  • Transcription Factor Research: Elucidates spatial arrangements of TF binding sites to support studies of gene regulatory mechanisms.
  • Genomic Annotation and Enhancer Analysis: Leverages genome annotation data to investigate enhancer regions and their potential TF occupancy patterns.

Methodology:

Analyzes genome annotation files using a cut-off distance parameter to define clusters, integrates position-weight matrix tools and genome annotation data from resources like PWMScan, compares observed clusters to simulated random distributions for statistical validation, and is implemented as an R package.

Topics

Details

Programming Languages:
R
Added:
1/9/2020
Last Updated:
12/10/2020

Operations

Publications

Golas S. CCSeq: Clusters of Colocalized Sequences. Unknown Journal. 2019. doi:10.1101/818385.