LDA-DC

LDA-DC performs probabilistic double clustering to jointly stratify patients and cell populations from high-dimensional omics datasets such as flow cytometry and microbiota profiles.


Key Features:

  • Double clustering: Simultaneously identifies clusters of patients and related clusters of cells within a unified model.
  • LDA extension: Implements an adaptation of Latent Dirichlet Allocation (LDA) to model biological observations rather than text topics.
  • Probabilistic generative model: Uses a generative probabilistic framework to represent latent patient- and cell-level structure.
  • Joint phenotype extraction: Extracts meaningful phenotypes from both patient-level and cell-level data concurrently.
  • Handles cohort variability: Designed to account for cohort variability and patient heterogeneity in high-dimensional datasets.
  • Validation on synthetic data: Validated using artificial datasets to assess performance on heterogeneous, large-scale data.
  • Application to cytometry and microbiota data: Applied to flow cytometry and microbiota datasets for real-world stratification tasks.
  • Computational efficiency: Described as computationally efficient for handling large omics datasets.

Scientific Applications:

  • Patient stratification: Stratifies patient cohorts into distinct subgroups based on multi-level omics signals.
  • Diagnostics: Supports diagnostic inference by linking patient clusters to characteristic cell-population profiles.
  • Identification of disease-associated cell populations: Discovers cell clusters associated with specific patient conditions.
  • Analysis of flow cytometry and microbiota datasets: Enables joint analysis of cytometry and microbiota data for integrative studies.
  • Pre-clinical research and clinical applications: Facilitates investigation of disease mechanisms and patient-specific conditions in pre-clinical and clinical contexts.

Methodology:

An extension of Latent Dirichlet Allocation (LDA) is used as a probabilistic generative model for double clustering, validated on artificial datasets and applied to flow cytometry and microbiota data to jointly infer patient clusters and associated cell-population clusters.

Topics

Details

License:
Not licensed
Cost:
Free of charge
Tool Type:
command-line tool
Operating Systems:
Mac, Linux, Windows
Programming Languages:
Python
Added:
3/20/2023
Last Updated:
11/24/2024

Operations

Publications

El Hachem E, Sokolovska N, Soula H. Latent dirichlet allocation for double clustering (LDA-DC): discovering patients phenotypes and cell populations within a single Bayesian framework. BMC Bioinformatics. 2023;24(1). doi:10.1186/s12859-023-05177-4. PMID:36823548. PMCID:PMC9948385.