Structure
Structure performs model-based Bayesian clustering to infer population structure from multi-locus genotype data, assign individuals to populations, estimate allele frequencies, and detect admixture and migrants while analyzing linkage disequilibrium patterns.
Key Features:
- Model-Based Clustering: Implements a probabilistic Bayesian framework that assumes an unknown number K of populations and estimates allele frequencies to assign individuals to one or more populations.
- Handling Linkage Disequilibrium: Accounts for linkage disequilibrium between loci to detect historical admixture, infer origins of chromosomal regions, and improve uncertainty estimates when loci are linked.
- Adaptability to Various Genetic Markers: Supports microsatellites and amplified fragment length polymorphisms (AFLPs) without assuming specific mutation processes.
- Handling Genotypic Ambiguity: Includes methods for dominant markers (e.g., AFLPs) to manage genotypic ambiguity, null alleles, and genotype-calling limitations in polyploids.
- Incorporation of Sample Group Information: Integrates sample group information by adjusting prior distributions for population assignments based on location-specific proportions.
- Cluster Number Estimation: Uses the log probability of the data alongside the ad hoc DeltaK statistic to assist estimation of the number of clusters K and detect hierarchical structure.
Scientific Applications:
- Population Structure Analysis: Infers and characterizes population structure from genotype data across species.
- Admixture Studies: Identifies admixed individuals and historical admixture events for studies of hybrid zones and migration.
- Genetic Diversity Surveys: Surveys genetic variation across multiple loci for conservation genetics and evolutionary analyses.
- Population Subdivision Detection: Detects subtle population subdivisions that may be missed by other methods.
Methodology:
Uses a probabilistic Bayesian clustering framework that assumes an unknown K and allele-frequency parameters, accounts for linkage disequilibrium, incorporates sample group priors, employs the DeltaK ad hoc statistic for cluster-number inference, and iteratively refines population assignments.
Topics
Collections
Details
- License:
- Not licensed
- Tool Type:
- command-line tool
- Operating Systems:
- Linux
- Programming Languages:
- Java, C
- Added:
- 8/20/2017
- Last Updated:
- 11/24/2024
Operations
Publications
Pritchard JK, Stephens M, Donnelly P. Inference of Population Structure Using Multilocus Genotype Data. Genetics. 2000;155(2):945-959. doi:10.1093/genetics/155.2.945. PMID:10835412. PMCID:PMC1461096.
Falush D, Stephens M, Pritchard JK. Inference of Population Structure Using Multilocus Genotype Data: Linked Loci and Correlated Allele Frequencies. Genetics. 2003;164(4):1567-1587. doi:10.1093/genetics/164.4.1567. PMID:12930761. PMCID:PMC1462648.
FALUSH D, STEPHENS M, PRITCHARD JK. Inference of population structure using multilocus genotype data: dominant markers and null alleles. Molecular Ecology Notes. 2007;7(4):574-578. doi:10.1111/j.1471-8286.2007.01758.x. PMID:18784791. PMCID:PMC1974779.
HUBISZ MJ, FALUSH D, STEPHENS M, PRITCHARD JK. Inferring weak population structure with the assistance of sample group information. Molecular Ecology Resources. 2009;9(5):1322-1332. doi:10.1111/j.1755-0998.2009.02591.x. PMID:21564903. PMCID:PMC3518025.
EVANNO G, REGNAUT S, GOUDET J. Detecting the number of clusters of individuals using the software <scp>structure</scp>: a simulation study. Molecular Ecology. 2005;14(8):2611-2620. doi:10.1111/j.1365-294x.2005.02553.x. PMID:15969739.