CAMLU

CAMLU employs an autoencoder-based supervised learning framework to annotate cell types in single-cell RNA sequencing (scRNA-seq) datasets and to detect novel cell types via reconstruction errors.


Key Features:

  • Autoencoder Integration: Employs an autoencoder trained on labeled data to reconstruct input sequences and compute reconstruction errors on testing datasets.
  • Iterative Feature Selection: Implements iterative feature selection that identifies features exhibiting a bi-modal pattern in their distribution to refine analyses.
  • Novel Cell Type Identification: Uses autoencoder reconstruction errors to pinpoint cells absent from the training set, reducing the number of unlabeled cells.
  • Support Vector Machine (SVM) Integration: Integrates a support vector machine for classification and annotation across detected cell types.
  • Validation on Real scRNA-seq Data: Demonstrated performance through numerical experiments on five real scRNA-seq datasets showing superior results compared to existing methods.

Scientific Applications:

  • Developmental Biology: Enables precise annotation of cell types to study developmental trajectories and tissue differentiation.
  • Immunology: Facilitates identification of known and novel immune cell subtypes in heterogeneous immune populations.
  • Oncology: Supports characterization of tumor heterogeneity by detecting distinct and novel tumor-associated cell types.

Methodology:

Trains an autoencoder on labeled scRNA-seq training data, reconstructs inputs and computes reconstruction errors on testing data, performs iterative feature selection by identifying bi-modal features, uses reconstruction errors to flag novel cells and reduce unlabeled cells, and integrates a support vector machine for classification.

Topics

Details

License:
CC-BY-4.0
Cost:
Free of charge
Tool Type:
library
Operating Systems:
Mac, Linux, Windows
Programming Languages:
R
Added:
10/28/2022
Last Updated:
11/24/2024

Operations

Publications

Li Z, Wang Y, Ganan-Gomez I, Colla S, Do K. A machine learning-based method for automatically identifying novel cells in annotating single-cell RNA-seq data. Bioinformatics. 2022;38(21):4885-4892. doi:10.1093/bioinformatics/btac617. PMID:36083008. PMCID:PMC9801963.

PMID: 36083008
PMCID: PMC9801963
Funding: - National Institutes of Health: 5R01GM122775, P30CA016672, P50CA140388, R03CA270725, UL1TR003167 - Cancer Prevention & Research Institute of Texas: RP160693 - Leukemia and Lymphoma Society. CPRIT: RP190295

Links