STARRPeaker
STARRPeaker performs statistical analysis of STARR-seq data to identify active regulatory regions (enhancers) while correcting for sequencing biases and overdispersion.
Key Features:
- Negative binomial regression framework: Implements a negative binomial regression model tailored for STARR-seq to model counts, mitigate overdispersion, and accommodate non-uniform coverage.
- Bias correction with covariates: Incorporates covariates such as GC content, mappability, and conservation to adjust for sequencing and genomic biases.
- Mitigation of confounders: Addresses confounding effects that arise in STARR-seq data, including biases related to GC content and library complexity from deep sequencing.
- Genome-wide enhancer identification: Applied to whole-genome STARR-seq datasets to identify enhancers across the genome, including analyses in human cell lines HepG2 and K562.
Scientific Applications:
- Enhancer discovery: Identification of active regulatory regions (enhancers) from STARR-seq experiments on a genome-wide scale.
- Gene regulation and chromatin studies: Facilitates interpretation of STARR-seq results for studies of gene regulation and chromatin dynamics by correcting for technical and genomic confounders.
Methodology:
Integration of covariates (GC content, mappability, conservation) into a negative binomial regression model to adjust for sequencing biases and overdispersion.
Topics
Details
- License:
- GPL-3.0
- Programming Languages:
- Python
- Added:
- 1/18/2021
- Last Updated:
- 2/24/2021
Operations
Publications
Lee D, Shi M, Moran J, Wall M, Zhang J, Liu J, Fitzgerald D, Kyono Y, Ma L, White KP, Gerstein M. STARRPeaker: uniform processing and accurate identification of STARR-seq active regions. Genome Biology. 2020;21(1). doi:10.1186/s13059-020-02194-x. PMID:33292397. PMCID:PMC7722316.
PMID: 33292397
PMCID: PMC7722316
Funding: - National Human Genome Research Institute: U24HG009446, UM1HG009426
- National Institute of Mental Health: K01MH123896