FSQN

FSQN performs feature-specific quantile normalization to remove platform-based technical bias and enable integration of RNA-seq with DNA microarray data for cross-platform gene expression analysis and molecular subtype classification.


Key Features:

  • Normalization Across Platforms: FSQN removes platform-based bias from RNA-seq data to allow consistent comparison across different gene expression profiling technologies.
  • R package implementation: FSQN is implemented as an R package.
  • Machine Learning Integration: The method employs machine learning classifiers trained on DNA microarray data to assign molecular subtypes to normalized RNA-seq datasets.
  • High Accuracy Rates: FSQN achieved up to 98% accuracy for breast invasive carcinoma (BRCA) and 97% accuracy for colorectal cancer (CRC) in subtype classification.
  • Versatility with Feature Scaling: The method maintains high performance regardless of feature scaling or the choice of machine learning algorithm.
  • Sample Size Considerations: Maximum accuracy is achieved when RNA-seq datasets contain at least 25 samples.

Scientific Applications:

  • Molecular subtype analysis in cancers: Integration of RNA-seq and DNA microarray data for subtype classification in breast invasive carcinoma (BRCA) and colorectal cancer (CRC).
  • Cross-platform studies of disease biology: Normalization and classification across platforms to study molecular heterogeneity, disease pathogenesis, and therapeutic responses in cancers and autoimmune diseases.

Methodology:

FSQN applies feature-specific quantile normalization and trains machine learning classifiers on DNA microarray data to normalize RNA-seq datasets and assign microarray-derived molecular subtypes.

Topics

Collections

Details

Tool Type:
library
Programming Languages:
R
Added:
1/20/2021
Last Updated:
5/17/2021

Operations

Publications

Franks JM, Cai G, Whitfield ML. Feature specific quantile normalization enables cross-platform classification of molecular subtypes using gene expression data. Bioinformatics. 2018;34(11):1868-1874. doi:10.1093/bioinformatics/bty026. PMID:29360996. PMCID:PMC5972664.

PMID: 29360996
PMCID: PMC5972664
Funding: - National Institutes of Health: P30 AR061271, P50 AR060780