subHMM

subHMM infers allele-specific copy number alterations (ASCNA), subclone genotypes, and clonal proportions from tumor genomic data using a hidden Markov model to characterize tumor subclonal architecture.


Key Features:

  • ASCNA and subclone region identification: Identifies genomic regions harboring distinct tumor subclones and allele-specific copy number alterations.
  • Genotype and clonal proportion estimation: Estimates subclone genotypes and their clonal proportions across the genome.
  • Two-step parameter estimation: Fits a standard hidden Markov model using a conglomerated hidden state for clonal genotype and subclone status, then derives region-specific clonal proportions by maximizing pseudo-likelihoods.
  • Statistical inference methods: Employs the Expectation-Maximization algorithm and the Forward-backward algorithm for parameter estimation and state decoding within the HMM framework.
  • Validation: Validated by simulation studies demonstrating accurate identification and characterization of subclone regions and genotypes in analyses of somatic copy number alterations.

Scientific Applications:

  • Tumor heterogeneity and evolution studies: Characterizes subclonal architecture to support analyses of tumor progression and heterogeneity.
  • TCGA renal cell carcinoma analyses: Applied to renal cell carcinoma datasets from The Cancer Genome Atlas (TCGA) to study subclones and somatic copy number alterations in real-world genomic data.

Methodology:

Fits a hidden Markov model with a hidden state variable representing combined clonal genotype and subclone status, uses the Expectation-Maximization algorithm and the Forward-backward algorithm for parameter estimation and state inference, and refines region-specific clonal proportions by maximizing pseudo-likelihoods.

Topics

Details

Programming Languages:
R
Added:
1/18/2021
Last Updated:
11/24/2024

Operations

Publications

Choo-Wosoba H, Albert PS, Zhu B. A hidden Markov modeling approach for identifying tumor subclones in next-generation sequencing studies. Biostatistics. 2020;23(1):69-82. doi:10.1093/biostatistics/kxaa013. PMID:32282873. PMCID:PMC9119345.