SPsimSeq
SPsimSeq generates simulated bulk and single-cell RNA sequencing (RNA-seq) data that reproduce gene expression marginal distributions and inter-gene dependencies for benchmarking and method evaluation.
Key Features:
- Semi-parametric density estimation: SPsimSeq uses a specially designed exponential family to estimate gene expression distributions from real bulk or single-cell RNA-seq data.
- Gaussian-copula dependence modeling: It employs Gaussian-copulas to simulate datasets from the estimated marginals while preserving inter-gene dependence structure.
- Modeling biological signals and confounders: The method can incorporate differential expression signals and account for confounding batch effects across experimental scenarios.
- Flexible experimental design simulation: It supports simulation of multiple groups and batches with specified sample sizes and library sizes.
- Implementation: SPsimSeq is implemented as an R package.
Scientific Applications:
- Benchmarking and method evaluation: Generating realistic RNA-seq datasets for validating bioinformatics pipelines and benchmarking analysis methods.
- Power analysis: Enabling power analyses by simulating datasets with controlled sample sizes, library sizes, and effect sizes.
- Method development and testing: Testing statistical models and algorithms under realistic dependence structures and marginal distributions.
- Batch effect assessment: Studying the impact of batch effects and other confounding factors on gene expression analyses.
Methodology:
SPsimSeq constructs marginal gene expression distributions from real RNA-seq data via a specially designed exponential-family density estimator and then generates simulated datasets by sampling from these marginals using Gaussian-copulas to preserve inter-gene dependencies.
Topics
Details
- Programming Languages:
- R
- Added:
- 1/18/2021
- Last Updated:
- 2/21/2021
Operations
Publications
Assefa AT, Vandesompele J, Thas O. SPsimSeq: semi-parametric simulation of bulk and single-cell RNA-sequencing data. Bioinformatics. 2020;36(10):3276-3278. doi:10.1093/bioinformatics/btaa105. PMID:32065619. PMCID:PMC7214028.
PMID: 32065619
PMCID: PMC7214028
Funding: - UGent Special Research Fund Concerted Research Actions: BOF16-GOA-023