Structure

Structure performs model-based Bayesian clustering to infer population structure from multi-locus genotype data, assign individuals to populations, estimate allele frequencies, and detect admixture and migrants while analyzing linkage disequilibrium patterns.


Key Features:

  • Model-Based Clustering: Implements a probabilistic Bayesian framework that assumes an unknown number K of populations and estimates allele frequencies to assign individuals to one or more populations.
  • Handling Linkage Disequilibrium: Accounts for linkage disequilibrium between loci to detect historical admixture, infer origins of chromosomal regions, and improve uncertainty estimates when loci are linked.
  • Adaptability to Various Genetic Markers: Supports microsatellites and amplified fragment length polymorphisms (AFLPs) without assuming specific mutation processes.
  • Handling Genotypic Ambiguity: Includes methods for dominant markers (e.g., AFLPs) to manage genotypic ambiguity, null alleles, and genotype-calling limitations in polyploids.
  • Incorporation of Sample Group Information: Integrates sample group information by adjusting prior distributions for population assignments based on location-specific proportions.
  • Cluster Number Estimation: Uses the log probability of the data alongside the ad hoc DeltaK statistic to assist estimation of the number of clusters K and detect hierarchical structure.

Scientific Applications:

  • Population Structure Analysis: Infers and characterizes population structure from genotype data across species.
  • Admixture Studies: Identifies admixed individuals and historical admixture events for studies of hybrid zones and migration.
  • Genetic Diversity Surveys: Surveys genetic variation across multiple loci for conservation genetics and evolutionary analyses.
  • Population Subdivision Detection: Detects subtle population subdivisions that may be missed by other methods.

Methodology:

Uses a probabilistic Bayesian clustering framework that assumes an unknown K and allele-frequency parameters, accounts for linkage disequilibrium, incorporates sample group priors, employs the DeltaK ad hoc statistic for cluster-number inference, and iteratively refines population assignments.

Topics

Collections

Details

License:
Not licensed
Tool Type:
command-line tool
Operating Systems:
Linux
Programming Languages:
Java, C
Added:
8/20/2017
Last Updated:
11/24/2024

Operations

Publications

Pritchard JK, Stephens M, Donnelly P. Inference of Population Structure Using Multilocus Genotype Data. Genetics. 2000;155(2):945-959. doi:10.1093/genetics/155.2.945. PMID:10835412. PMCID:PMC1461096.

Falush D, Stephens M, Pritchard JK. Inference of Population Structure Using Multilocus Genotype Data: Linked Loci and Correlated Allele Frequencies. Genetics. 2003;164(4):1567-1587. doi:10.1093/genetics/164.4.1567. PMID:12930761. PMCID:PMC1462648.

FALUSH D, STEPHENS M, PRITCHARD JK. Inference of population structure using multilocus genotype data: dominant markers and null alleles. Molecular Ecology Notes. 2007;7(4):574-578. doi:10.1111/j.1471-8286.2007.01758.x. PMID:18784791. PMCID:PMC1974779.

HUBISZ MJ, FALUSH D, STEPHENS M, PRITCHARD JK. Inferring weak population structure with the assistance of sample group information. Molecular Ecology Resources. 2009;9(5):1322-1332. doi:10.1111/j.1755-0998.2009.02591.x. PMID:21564903. PMCID:PMC3518025.

EVANNO G, REGNAUT S, GOUDET J. Detecting the number of clusters of individuals using the software <scp>structure</scp>: a simulation study. Molecular Ecology. 2005;14(8):2611-2620. doi:10.1111/j.1365-294x.2005.02553.x. PMID:15969739.

Documentation