pcqc

pcqc selects principal components that capture informative variance in single-cell RNA sequencing (scRNA-seq) data to improve clustering and the detection of rare cell populations.


Key Features:

  • Optimal Principal Component Selection: Uses a quantile-based criterion focusing on the tails of the per-observation distribution of variance explained to prioritize components that capture signals from rare cell populations.
  • Computational Efficiency: Offers a computationally efficient alternative to existing selection methods for large-scale scRNA-seq datasets.
  • Robustness Across Datasets: Evaluated across three single-cell RNA-sequencing datasets and shown to match or outperform traditional variance-explained selection criteria.

Scientific Applications:

  • Preprocessing for scRNA-seq: Selects principal components for downstream clustering and analysis in single-cell RNA-seq workflows.
  • Dimensionality Reduction and Visualization: Improves inputs for non-linear dimensionality reduction methods such as t-SNE and UMAP to yield more accurate representations of cell populations.
  • Rare Cell Population Detection: Enhances identification of rare cell subtypes that may be underrepresented by traditional variance-explained approaches.

Methodology:

pcqc applies a quantile-based approach to the distribution of variance explained per observation to identify and select principal components emphasizing signal in the distribution tails.

Topics

Details

Programming Languages:
Python
Added:
1/18/2021
Last Updated:
1/23/2021

Operations

Publications

Burstein D, Fullard JF, Roussos P. PCQC: Selecting optimal principal components for identifying clusters with highly imbalanced class sizes in single-cell RNA-seq data. Unknown Journal. 2020. doi:10.1101/2020.11.19.390542.