CAEN
CAEN selects feature genes for classifying bulk and single-cell RNA-seq data by using rank-based category encoding and correlation analysis to identify differentially expressed (DE) genes that distinguish classes.
Key Features:
- Category Encoding via Rank: Encodes categories using the rank of sequence samples for each gene across different classes to represent class-specific expression patterns.
- Correlation Coefficients Analysis: Computes correlation coefficients between genes and classes based on sample ranks and a defined category rank to quantify association strength.
- Feature Gene Identification: Selects genes with the highest correlation coefficients as feature genes for downstream classification and dimensionality reduction.
- Sure Screening Method: Applies a sure screening approach to ensure rank consistency properties among selected features.
- Nonparametric Design (vs BSS/WSS): Operates without assuming normal data distributions, diverging from BSS/WSS (BW) approaches that rely on parametric assumptions.
- Performance Evaluation: Demonstrates classifier performance through simulation studies and analyses of real datasets, showing comparable or superior results to competing methods in most settings.
Scientific Applications:
- RNA-seq classification: Improves classification of bulk and single-cell RNA-seq samples by identifying discriminative feature genes.
- Disease diagnosis: Enables selection of gene sets that can distinguish disease states from controls in expression-based diagnostic studies.
- Cellular heterogeneity analysis: Facilitates discovery of genes that separate cellular subpopulations in single-cell RNA-seq datasets.
- Molecular mechanism exploration: Supports investigation of genes underlying different biological conditions by highlighting DE feature genes relevant to class differences.
Methodology:
CAEN performs rank-based category encoding, computes correlation coefficients between genes and classes using sample ranks and a category rank, selects genes with highest correlations, and applies a sure screening step to ensure rank consistency while avoiding assumptions of normality (contrasting with BSS/WSS (BW)).
Topics
Details
- License:
- GPL-2.0
- Tool Type:
- library
- Programming Languages:
- R
- Added:
- 6/14/2021
- Last Updated:
- 8/18/2021
Operations
Publications
Zhou Y, Zhang L, Xu J, Zhang J, Yan X. Category encoding method to select feature genes for the classification of bulk and single‐cell RNA‐seq data. Statistics in Medicine. 2021;40(18):4077-4089. doi:10.1002/sim.9015. PMID:34028849.