SampleSeq2

SampleSeq2 optimizes sample selection for targeted resequencing by selecting optimally unrelated individuals to maximize representation of founder chromosomes (G(T)) and minimize identical-by-descent (IBD) redundancy.


Key Features:

  • Probability-based algorithm: Estimates the number of founder chromosomes denoted as G(T) within a sample using a probability-based approach.
  • Selection of optimally unrelated individuals: Identifies and selects individuals to minimize regions inherited identical by descent (IBD) from common ancestors.
  • Minimum sample-size determination: Determines the minimum number of subjects required to represent founder chromosomes effectively.
  • Cost-mitigation for sequencing studies: Facilitates strategic selection of subsets for targeted resequencing to reduce sampling redundancy relative to whole-genome or exome sequencing.
  • Enhanced association power: Increases the power of association tests and reduces the impact of cryptic relatedness on parameter estimates.
  • Allele yield and IBD reduction: Increases total yield of alleles from sequencing and minimizes the average size of IBD regions around disease-associated alleles in cases.
  • Empirical validation: Demonstrated increased G(T) compared to random sampling across small sample sizes in simulations and the Anabaptist genealogy dataset.

Scientific Applications:

  • Targeted resequencing study design: Selecting subsets of subjects from larger cohorts to maximize genetic diversity for targeted resequencing experiments.
  • Genome-wide agnostic investigations: Enabling agnostic investigations across the genome by minimizing redundant IBD regions among selected samples.
  • Association studies: Improving statistical power and robustness of association tests by reducing cryptic relatedness.
  • Founder and pedigree population analysis: Optimizing sample selection in founder or genealogy datasets such as the Anabaptist genealogy dataset to capture independent ancestral lineages.
  • Rare variant discovery and localization: Increasing allele yield and reducing shared IBD around disease alleles to aid discovery and localization of disease-associated variants.

Methodology:

Implements a probability-based algorithm to estimate founder chromosomes (G(T)) and compute minimum sample sizes, with validation via simulations and application to the Anabaptist genealogy dataset.

Topics

Details

Tool Type:
library
Operating Systems:
Linux, Windows, Mac
Programming Languages:
R, Perl
Added:
8/3/2017
Last Updated:
11/25/2024

Operations

Publications

Edwards TL, Li C. Optimized Selection of Unrelated Subjects for Whole‐Genome Sequencing Studies of Rare High‐Penetrance Alleles. Genetic Epidemiology. 2012;36(5):472-479. doi:10.1002/gepi.21641. PMID:22623060. PMCID:PMC3738264.

Documentation

Links