SampleSeq2
SampleSeq2 optimizes sample selection for targeted resequencing by selecting optimally unrelated individuals to maximize representation of founder chromosomes (G(T)) and minimize identical-by-descent (IBD) redundancy.
Key Features:
- Probability-based algorithm: Estimates the number of founder chromosomes denoted as G(T) within a sample using a probability-based approach.
- Selection of optimally unrelated individuals: Identifies and selects individuals to minimize regions inherited identical by descent (IBD) from common ancestors.
- Minimum sample-size determination: Determines the minimum number of subjects required to represent founder chromosomes effectively.
- Cost-mitigation for sequencing studies: Facilitates strategic selection of subsets for targeted resequencing to reduce sampling redundancy relative to whole-genome or exome sequencing.
- Enhanced association power: Increases the power of association tests and reduces the impact of cryptic relatedness on parameter estimates.
- Allele yield and IBD reduction: Increases total yield of alleles from sequencing and minimizes the average size of IBD regions around disease-associated alleles in cases.
- Empirical validation: Demonstrated increased G(T) compared to random sampling across small sample sizes in simulations and the Anabaptist genealogy dataset.
Scientific Applications:
- Targeted resequencing study design: Selecting subsets of subjects from larger cohorts to maximize genetic diversity for targeted resequencing experiments.
- Genome-wide agnostic investigations: Enabling agnostic investigations across the genome by minimizing redundant IBD regions among selected samples.
- Association studies: Improving statistical power and robustness of association tests by reducing cryptic relatedness.
- Founder and pedigree population analysis: Optimizing sample selection in founder or genealogy datasets such as the Anabaptist genealogy dataset to capture independent ancestral lineages.
- Rare variant discovery and localization: Increasing allele yield and reducing shared IBD around disease alleles to aid discovery and localization of disease-associated variants.
Methodology:
Implements a probability-based algorithm to estimate founder chromosomes (G(T)) and compute minimum sample sizes, with validation via simulations and application to the Anabaptist genealogy dataset.
Topics
Details
- Tool Type:
- library
- Operating Systems:
- Linux, Windows, Mac
- Programming Languages:
- R, Perl
- Added:
- 8/3/2017
- Last Updated:
- 11/25/2024
Operations
Publications
Edwards TL, Li C. Optimized Selection of Unrelated Subjects for Whole‐Genome Sequencing Studies of Rare High‐Penetrance Alleles. Genetic Epidemiology. 2012;36(5):472-479. doi:10.1002/gepi.21641. PMID:22623060. PMCID:PMC3738264.