scSemiCluster

scSemiCluster performs semi-supervised clustering and annotation of single-cell RNA-seq data by integrating labeled reference and unlabeled target datasets to improve cell-type identification.


Key Features:

  • Integration of Reference and Target Data: Integrates labeled reference and unlabeled target scRNA-seq datasets during training to leverage information from both domains.
  • Structure Similarity Regularization: Applies structure similarity regularization within the reference domain to constrain clustering solutions in the target domain.
  • Pairwise Constraints in Feature Learning: Incorporates pairwise constraints so that cells from the same cluster are proximal and cells from different clusters are separated in the latent space.
  • Domain Adaptation without Explicit Alignment: Implements unsupervised domain adaptation principles to transfer labels without performing explicit domain alignment or batch effect correction.
  • Deep Clustering Techniques: Combines deep discriminative clustering and deep generative clustering approaches for latent representation and cluster definition.

Scientific Applications:

  • Annotation of Unlabeled scRNA-seq Datasets: Enables annotation of target single-cell RNA-seq datasets using information from labeled reference datasets.
  • Cell Type and Regulatory Gene Identification: Facilitates identification of cell types and regulatory genes through improved clustering and downstream analysis.
  • Biomedical Research Domains: Applicable to studies in developmental biology, immunology, and cancer research where cell-type resolution is required.

Methodology:

Integrates labeled reference and unlabeled target datasets in a semi-supervised framework, applies structure similarity regularization in the reference domain, enforces pairwise constraints during feature learning, combines deep discriminative and deep generative clustering, and follows unsupervised domain adaptation principles without explicit domain alignment or batch effect correction.

Topics

Details

Programming Languages:
Python
Added:
1/18/2021
Last Updated:
2/13/2021

Operations

Publications

Chen L, He Q, Zhai Y, Deng M. Single-cell RNA-seq data semi-supervised clustering and annotation via structural regularized domain adaptation. Bioinformatics. 2020;37(6):775-784. doi:10.1093/bioinformatics/btaa908. PMID:33098418.

PMID: 33098418
Funding: - National Key Research and Development Program of China: 2016YFA0502303 - National Key Basic Research Project of China: 2015CB910303 - National Natural Science Foundation of China: 31871342