SCHEMA

SCHEMA integrates heterogeneous data modalities from multi-modal single-cell datasets to synthesize information across gene expression, chromatin accessibility, spliced and unspliced mRNA, and peptide sequence features for downstream analyses.


Key Features:

  • Principled Metric Learning Strategy: SCHEMA employs a principled metric learning approach to synthesize multimodal information and capture complex relationships between modalities.
  • Flexibility and Power: SCHEMA accommodates an arbitrary number of data modalities, including gene expression, chromatin accessibility, and spliced/unspliced mRNA.
  • Computational Efficiency: SCHEMA is designed for computational efficiency on large-scale single-cell datasets.

Scientific Applications:

  • Cell Type Inference: Integrates gene expression with chromatin accessibility to infer cell types within a sample.
  • Differential Gene Expression Analysis: Performs differential gene expression analysis while accounting for confounders such as batch effects and developmental age.
  • Evolutionary Pressure Estimation: Estimates evolutionary pressures on peptide sequences.
  • Cell Differentiation Inference: Synthesizes spliced and unspliced mRNA data to infer cell differentiation processes.

Methodology:

Principled metric learning applied to simultaneous multi-modal single-cell measurements within a unified analytical framework.

Topics

Details

Programming Languages:
Python
Added:
1/14/2020
Last Updated:
1/16/2021

Operations

Publications

Singh R, Hie B, Narayan A, Berger B. Metric learning enables synthesis of heterogeneous single-cell modalities. Unknown Journal. 2019. doi:10.1101/834549.