subHMM
subHMM infers allele-specific copy number alterations (ASCNA), subclone genotypes, and clonal proportions from tumor genomic data using a hidden Markov model to characterize tumor subclonal architecture.
Key Features:
- ASCNA and subclone region identification: Identifies genomic regions harboring distinct tumor subclones and allele-specific copy number alterations.
- Genotype and clonal proportion estimation: Estimates subclone genotypes and their clonal proportions across the genome.
- Two-step parameter estimation: Fits a standard hidden Markov model using a conglomerated hidden state for clonal genotype and subclone status, then derives region-specific clonal proportions by maximizing pseudo-likelihoods.
- Statistical inference methods: Employs the Expectation-Maximization algorithm and the Forward-backward algorithm for parameter estimation and state decoding within the HMM framework.
- Validation: Validated by simulation studies demonstrating accurate identification and characterization of subclone regions and genotypes in analyses of somatic copy number alterations.
Scientific Applications:
- Tumor heterogeneity and evolution studies: Characterizes subclonal architecture to support analyses of tumor progression and heterogeneity.
- TCGA renal cell carcinoma analyses: Applied to renal cell carcinoma datasets from The Cancer Genome Atlas (TCGA) to study subclones and somatic copy number alterations in real-world genomic data.
Methodology:
Fits a hidden Markov model with a hidden state variable representing combined clonal genotype and subclone status, uses the Expectation-Maximization algorithm and the Forward-backward algorithm for parameter estimation and state inference, and refines region-specific clonal proportions by maximizing pseudo-likelihoods.
Topics
Details
- Programming Languages:
- R
- Added:
- 1/18/2021
- Last Updated:
- 11/24/2024
Operations
Publications
Choo-Wosoba H, Albert PS, Zhu B. A hidden Markov modeling approach for identifying tumor subclones in next-generation sequencing studies. Biostatistics. 2020;23(1):69-82. doi:10.1093/biostatistics/kxaa013. PMID:32282873. PMCID:PMC9119345.