DiviK

DiviK segments high-dimensional datasets such as mass spectrometry imaging (MSI) and other omics data to reveal molecular heterogeneity and spatial tissue substructures.


Key Features:

  • Scalability: Processes datasets exceeding 1.5 million instances with thousands of features, enabling large-scale omics analyses including single-cell spatial transcriptomics and embedded medical imaging.
  • Stepwise deglomerative algorithm: Employs a scalable stepwise deglomerative schema that iteratively refines clusters to uncover intricate data patterns.
  • Local data-driven feature space adaptation: Uses Gaussian Mixture Model (GMM)-based local optimization to adapt the feature space dynamically and reduce dependence on global feature engineering.
  • Feature engineering compatibility: Evaluated with feature engineering approaches including None, PCA, EXIMS, UMAP, and Neural Ions to assess impact on clustering quality.
  • Algorithmic simplicity and flexibility: Implements a simple, generalizable algorithm applicable to MSI, other omics datasets, and tabular data formats.

Scientific Applications:

  • Tumor analysis: Segments high-dimensional MSI datasets from 2D human cancer tissue samples to characterize tumor origin and composition.
  • Metabolomics studies: Identifies spatially consistent clusters and overall tissue composition relevant to metabolic investigations.
  • 3D organ imaging: Applied to 3D mouse kidney images for spatial heterogeneity analysis.
  • Single-cell and medical imaging: Applicable to single-cell spatial transcriptomics and to medical imaging after appropriate embedding.

Methodology:

Uses a scalable stepwise deglomerative algorithm with Gaussian Mixture Model (GMM)-based local optimization; compares to clustering algorithms including regular k-means, spatial, and spectral approaches; evaluates feature engineering methods (None, PCA, EXIMS, UMAP, Neural Ions) and measures performance with Dice Index, Rand Index, and EXIMS score focusing on clustering composition, tumor region coverage, and spatial consistency.

Topics

Details

License:
Apache-2.0
Cost:
Free of charge
Tool Type:
library
Operating Systems:
Mac, Linux, Windows
Programming Languages:
Python
Added:
2/19/2023
Last Updated:
11/24/2024

Operations

Publications

Mrukwa G, Polanska J. DiviK: divisive intelligent K-means for hands-free unsupervised clustering in big biological data. BMC Bioinformatics. 2022;23(1). doi:10.1186/s12859-022-05093-z. PMID:36503372. PMCID:PMC9743550.

PMID: 36503372
PMCID: PMC9743550
Funding: - Narodowe Centrum Nauki: UMO-2015/19/B/ST6/01736 - Narodowe Centrum Badań i Rozwoju: I029/17-POWR.03.02.00-IP.08-00-DOK/17

Downloads

Links