SCHEMA
SCHEMA integrates heterogeneous data modalities from multi-modal single-cell datasets to synthesize information across gene expression, chromatin accessibility, spliced and unspliced mRNA, and peptide sequence features for downstream analyses.
Key Features:
- Principled Metric Learning Strategy: SCHEMA employs a principled metric learning approach to synthesize multimodal information and capture complex relationships between modalities.
- Flexibility and Power: SCHEMA accommodates an arbitrary number of data modalities, including gene expression, chromatin accessibility, and spliced/unspliced mRNA.
- Computational Efficiency: SCHEMA is designed for computational efficiency on large-scale single-cell datasets.
Scientific Applications:
- Cell Type Inference: Integrates gene expression with chromatin accessibility to infer cell types within a sample.
- Differential Gene Expression Analysis: Performs differential gene expression analysis while accounting for confounders such as batch effects and developmental age.
- Evolutionary Pressure Estimation: Estimates evolutionary pressures on peptide sequences.
- Cell Differentiation Inference: Synthesizes spliced and unspliced mRNA data to infer cell differentiation processes.
Methodology:
Principled metric learning applied to simultaneous multi-modal single-cell measurements within a unified analytical framework.
Topics
Details
- Programming Languages:
- Python
- Added:
- 1/14/2020
- Last Updated:
- 1/16/2021
Operations
Publications
Singh R, Hie B, Narayan A, Berger B. Metric learning enables synthesis of heterogeneous single-cell modalities. Unknown Journal. 2019. doi:10.1101/834549.
DOI: 10.1101/834549