MACML
MACML detects and characterizes clusters of specified site types (e.g., polymorphic sites or substituted sites) within DNA or protein sequences to profile sequence heterogeneity for studies of genetic variation and evolution.
Key Features:
- Clustering without prior knowledge: Identifies heterogeneous regions and clusters of specified site types without requiring prior specification of cluster size or number.
- Maximum likelihood hierarchical clustering: Uses a maximum likelihood framework for hierarchical clustering to model site-wise heterogeneity.
- Tripartite divide-and-conquer strategy: Implements a tripartite divide-and-conquer approach to partition the problem and model sequence complexity.
- Model evaluation and selection: Supports model comparison using Akaike Information Criterion (AIC), corrected AIC (AICc), and Bayesian Information Criterion (BIC).
- Model averaging: Produces averaged heterogeneity profiles by weighting models with their likelihoods to account for model uncertainty.
- Performance evaluation: Demonstrates greater detection power and improved accuracy and precision in simulations and empirical polymorphism data (including Drosophila alcohol dehydrogenase alleles) versus methods based on empirical cumulative distribution function statistics.
Scientific Applications:
- Detection of clustered polymorphic or substituted sites: Identifies and delineates clusters of variant sites within DNA and protein sequences.
- Genetic variation and evolutionary analysis: Profiles site-wise heterogeneity to inform studies of selection, mutation clustering, and evolutionary processes.
- Molecular genetics of specific loci: Applied to empirical polymorphism datasets such as natural alleles of the Drosophila alcohol dehydrogenase gene to characterize clustered variation.
Methodology:
Computational methods explicitly include maximum likelihood hierarchical clustering, a tripartite divide-and-conquer strategy, model selection via AIC/AICc/BIC, model averaging using weighted model likelihoods, and evaluation on simulated datasets and empirical polymorphism data (including Drosophila ADH) with comparisons to empirical cumulative distribution function–based methods.
Topics
Details
- Tool Type:
- command-line tool
- Operating Systems:
- Linux, Windows, Mac
- Programming Languages:
- C++
- Added:
- 12/18/2017
- Last Updated:
- 12/10/2018
Operations
Publications
Zhang Z, Townsend JP. Maximum-Likelihood Model Averaging To Profile Clustering of Site Types across Discrete Linear Sequences. PLoS Computational Biology. 2009;5(6):e1000421. doi:10.1371/journal.pcbi.1000421. PMID:19557160. PMCID:PMC2695770.