SPsimSeq

SPsimSeq generates simulated bulk and single-cell RNA sequencing (RNA-seq) data that reproduce gene expression marginal distributions and inter-gene dependencies for benchmarking and method evaluation.


Key Features:

  • Semi-parametric density estimation: SPsimSeq uses a specially designed exponential family to estimate gene expression distributions from real bulk or single-cell RNA-seq data.
  • Gaussian-copula dependence modeling: It employs Gaussian-copulas to simulate datasets from the estimated marginals while preserving inter-gene dependence structure.
  • Modeling biological signals and confounders: The method can incorporate differential expression signals and account for confounding batch effects across experimental scenarios.
  • Flexible experimental design simulation: It supports simulation of multiple groups and batches with specified sample sizes and library sizes.
  • Implementation: SPsimSeq is implemented as an R package.

Scientific Applications:

  • Benchmarking and method evaluation: Generating realistic RNA-seq datasets for validating bioinformatics pipelines and benchmarking analysis methods.
  • Power analysis: Enabling power analyses by simulating datasets with controlled sample sizes, library sizes, and effect sizes.
  • Method development and testing: Testing statistical models and algorithms under realistic dependence structures and marginal distributions.
  • Batch effect assessment: Studying the impact of batch effects and other confounding factors on gene expression analyses.

Methodology:

SPsimSeq constructs marginal gene expression distributions from real RNA-seq data via a specially designed exponential-family density estimator and then generates simulated datasets by sampling from these marginals using Gaussian-copulas to preserve inter-gene dependencies.

Topics

Details

Programming Languages:
R
Added:
1/18/2021
Last Updated:
2/21/2021

Operations

Publications

Assefa AT, Vandesompele J, Thas O. SPsimSeq: semi-parametric simulation of bulk and single-cell RNA-sequencing data. Bioinformatics. 2020;36(10):3276-3278. doi:10.1093/bioinformatics/btaa105. PMID:32065619. PMCID:PMC7214028.

PMID: 32065619
PMCID: PMC7214028
Funding: - UGent Special Research Fund Concerted Research Actions: BOF16-GOA-023