FSQN
FSQN performs feature-specific quantile normalization to remove platform-based technical bias and enable integration of RNA-seq with DNA microarray data for cross-platform gene expression analysis and molecular subtype classification.
Key Features:
- Normalization Across Platforms: FSQN removes platform-based bias from RNA-seq data to allow consistent comparison across different gene expression profiling technologies.
- R package implementation: FSQN is implemented as an R package.
- Machine Learning Integration: The method employs machine learning classifiers trained on DNA microarray data to assign molecular subtypes to normalized RNA-seq datasets.
- High Accuracy Rates: FSQN achieved up to 98% accuracy for breast invasive carcinoma (BRCA) and 97% accuracy for colorectal cancer (CRC) in subtype classification.
- Versatility with Feature Scaling: The method maintains high performance regardless of feature scaling or the choice of machine learning algorithm.
- Sample Size Considerations: Maximum accuracy is achieved when RNA-seq datasets contain at least 25 samples.
Scientific Applications:
- Molecular subtype analysis in cancers: Integration of RNA-seq and DNA microarray data for subtype classification in breast invasive carcinoma (BRCA) and colorectal cancer (CRC).
- Cross-platform studies of disease biology: Normalization and classification across platforms to study molecular heterogeneity, disease pathogenesis, and therapeutic responses in cancers and autoimmune diseases.
Methodology:
FSQN applies feature-specific quantile normalization and trains machine learning classifiers on DNA microarray data to normalize RNA-seq datasets and assign microarray-derived molecular subtypes.
Topics
Collections
Details
- Tool Type:
- library
- Programming Languages:
- R
- Added:
- 1/20/2021
- Last Updated:
- 5/17/2021
Operations
Publications
Franks JM, Cai G, Whitfield ML. Feature specific quantile normalization enables cross-platform classification of molecular subtypes using gene expression data. Bioinformatics. 2018;34(11):1868-1874. doi:10.1093/bioinformatics/bty026. PMID:29360996. PMCID:PMC5972664.