STARRPeaker

STARRPeaker performs statistical analysis of STARR-seq data to identify active regulatory regions (enhancers) while correcting for sequencing biases and overdispersion.


Key Features:

  • Negative binomial regression framework: Implements a negative binomial regression model tailored for STARR-seq to model counts, mitigate overdispersion, and accommodate non-uniform coverage.
  • Bias correction with covariates: Incorporates covariates such as GC content, mappability, and conservation to adjust for sequencing and genomic biases.
  • Mitigation of confounders: Addresses confounding effects that arise in STARR-seq data, including biases related to GC content and library complexity from deep sequencing.
  • Genome-wide enhancer identification: Applied to whole-genome STARR-seq datasets to identify enhancers across the genome, including analyses in human cell lines HepG2 and K562.

Scientific Applications:

  • Enhancer discovery: Identification of active regulatory regions (enhancers) from STARR-seq experiments on a genome-wide scale.
  • Gene regulation and chromatin studies: Facilitates interpretation of STARR-seq results for studies of gene regulation and chromatin dynamics by correcting for technical and genomic confounders.

Methodology:

Integration of covariates (GC content, mappability, conservation) into a negative binomial regression model to adjust for sequencing biases and overdispersion.

Topics

Details

License:
GPL-3.0
Programming Languages:
Python
Added:
1/18/2021
Last Updated:
2/24/2021

Operations

Publications

Lee D, Shi M, Moran J, Wall M, Zhang J, Liu J, Fitzgerald D, Kyono Y, Ma L, White KP, Gerstein M. STARRPeaker: uniform processing and accurate identification of STARR-seq active regions. Genome Biology. 2020;21(1). doi:10.1186/s13059-020-02194-x. PMID:33292397. PMCID:PMC7722316.

PMID: 33292397
PMCID: PMC7722316
Funding: - National Human Genome Research Institute: U24HG009446, UM1HG009426 - National Institute of Mental Health: K01MH123896