Cookie

Cookie selects representative samples from large single-cell sequencing datasets to enable experimental characterization of diverse cellular populations.


Key Features:

  • Relationship quantification: Vectorizes sample properties and computes Manhattan distances to quantify relationships and similarities among samples.
  • Sample size determination: Evaluates coverage of key properties across candidate sample sizes to determine an appropriate subset size that balances diverse properties and priority levels.
  • K-medoids clustering: Applies k-medoids clustering and selects cluster medoids as representative samples, leveraging k-medoids' robustness to noise and outliers.
  • Benchmarking across datasets: Demonstrates efficacy, efficiency, and flexibility via comparisons with conventional sampling methods on a single-cell atlas dataset, epidemiology surveillance data, and simulated datasets.
  • Implementation: Implemented in R.

Scientific Applications:

  • Single-cell sequencing sample selection: Selects representative subsets from high-dimensional single-cell sequencing data to support downstream experimental characterization.
  • Epidemiological surveillance sampling: Identifies representative samples from epidemiology surveillance datasets for monitoring and analysis.
  • Method benchmarking and simulation studies: Uses simulated datasets to evaluate sampling performance and compare against conventional methods.

Methodology:

Vectorize sample properties, compute Manhattan distances to quantify similarities, evaluate coverage across candidate sample sizes to choose an appropriate size, and perform k-medoids clustering to select cluster medoids as representative samples.

Topics

Details

License:
Not licensed
Cost:
Free of charge
Tool Type:
library
Operating Systems:
Mac, Linux, Windows
Programming Languages:
R
Added:
10/7/2022
Last Updated:
11/24/2024

Operations

Data Inputs & Outputs

Clustering

Outputs

    Publications

    Li L, Lan LY, Huang L, Ye C, Andrade J, Wilson PC. Selecting Representative Samples From Complex Biological Datasets Using K-Medoids Clustering. Frontiers in Genetics. 2022;13. doi:10.3389/fgene.2022.954024. PMID:35910222. PMCID:PMC9335369.

    PMID: 35910222
    PMCID: PMC9335369
    Funding: - National Institutes of Health: 2P01AI097092-06A1 U19AI109946 U19AI057266

    Documentation

    Links