TWO-SIGMA

TWO-SIGMA models differential expression and gene set testing for single-cell RNA sequencing (scRNA-seq) data using a two-component mixed-effects framework that accounts for dropout probability and conditional mean expression to handle zero-inflation, overdispersion, and within-sample correlation.


Key Features:

  • Two-Component Model: Implements a two-component framework that models drop-out probability with a mixed-effects logistic regression and conditional mean expression with a mixed-effects negative binomial regression.
  • Zero-Inflation and Overdispersion Handling: Explicitly models zero-inflation and overdispersion typical of scRNA-seq count data via the two-component specification and negative binomial variance structure.
  • Correlation Structure Accommodation: Incorporates random effect terms to account for correlations between cells from the same individual.
  • No Log-Transformation Required: Operates on count-scale outcomes without requiring log-transformation of expression data.
  • Unbalanced Design Support: Handles unbalanced designs with varying numbers of cells per sample.
  • Covariate Adjustment: Allows inclusion of covariates at both sample and cell levels, including batch effects.
  • Interpretable Effect Estimates: Produces interpretable effect size estimates and supports general differential expression tests beyond two-group comparisons.
  • Performance: Simulation studies demonstrate improved control of type-I error rates and enhanced statistical power relative to alternative regression-based approaches, particularly with moderate within-sample correlations.

Scientific Applications:

  • Developmental biology: Enables detection of cell-type–specific differential expression and gene set changes during development at single-cell resolution.
  • Cancer genomics: Supports analysis of tumor heterogeneity and differential expression across malignant and microenvironment cell populations in scRNA-seq datasets.
  • Immunology: Facilitates identification of immune cell subtype–specific expression changes and pathway activity in single-cell immune profiling.
  • Pancreas islet single-cell analysis: Applicable to pancreas islet single-cell datasets for evaluating gene expression differences while accounting for within-sample correlation.

Methodology:

TWO-SIGMA fits mixed-effects logistic regression for dropout probability and mixed-effects negative binomial regression for conditional mean expression, integrating both components and including random effect terms to model within-sample correlation.

Topics

Details

Programming Languages:
R
Added:
1/18/2021
Last Updated:
11/24/2024

Operations

Publications

Van Buren E, Hu M, Weng C, Jin F, Li Y, Wu D, Li Y. TWO‐SIGMA: A novel two‐component single cell model‐based association method for single‐cell RNA‐seq data. Genetic Epidemiology. 2020;45(2):142-153. doi:10.1002/gepi.22361. PMID:32989764. PMCID:PMC8570615.

PMID: 32989764
PMCID: PMC8570615
Funding: - National Institute of Dental and Craniofacial Research: R03DE028983 - National Institute of Diabetes and Digestive and Kidney Diseases: R01DK113185, U54DK107977 - National Human Genome Research Institute: R01HG009658 - National Heart, Lung, and Blood Institute: R01HL129132 - Eunice Kennedy Shriver National Institute of Child Health and Human Development: U54HD079124 - National Institute of General Medical Sciences: R01GM105785