UT-AIM250
UT-AIM250 estimates individual genetic ancestry proportions from whole exome sequencing (WES) data to improve three-way admixture inference in Hispanic and admixed populations.
Key Features:
- Ancestry Informative Marker Panel: Utilizes a panel of 250 single nucleotide polymorphisms (SNPs) selected from evolutionarily conserved exonic regions typically covered by WES.
- Three-Way Genetic Admixture Inference: Infers admixture proportions among African (AFR), European (EUR), and East Asian (EAS) ancestral populations.
- High Accuracy: Reported accuracies of 0.995 ± 0.012 for AFR, 0.997 ± 0.007 for EUR, and 0.994 ± 0.012 for EAS when tested on three ancestral populations.
- Application to Admixed Populations: Validated on admixed American (AMR) samples from the 1000 Genomes Project with performance comparable to previously published AIM panels.
- Clinical Application: Applied to a clinical cohort of 26 self-reported Hispanic hepatocellular carcinoma (HCC) patients in South Texas using WES from adjacent non-cancer liver tissues and corresponding tumor tissues.
- Consistency Across Tissues: Maintains consistent admixture estimates across tumor and matched normal tissues in The Cancer Genome Atlas (TCGA) hepatocellular carcinoma samples.
- Identification of Misreported Ethnicity: Can identify discrepancies between reported race/ethnicity and genetic ancestry estimates.
Scientific Applications:
- Sample Stratification: Controlling sample collection and patient stratification in genetic studies of admixed populations using WES-derived ancestry estimates.
- Genetic Association and Health Disparities: Reducing bias in genetic association studies of cancer susceptibility and other health disparities among Hispanic populations by providing precise ancestry estimates.
- Clinical and Tumor-Normal Analyses: Enabling ancestry-consistent analyses in clinical WES datasets and tumor-versus-normal comparisons, as demonstrated in the South Texas HCC cohort and TCGA.
Methodology:
Selected approximately 1 million exonic SNPs from African (AFR), European (EUR), and East Asian (EAS) populations in the 1000 Genomes Project and refined SNPs based on linkage disequilibrium and ancestral informativeness statistics to derive the 250-SNP panel optimized for ancestry estimation.
Topics
Details
- Programming Languages:
- R
- Added:
- 1/18/2021
- Last Updated:
- 3/10/2021
Operations
Publications
Wang L, Zhang CW, Su SC, Chen HH, Chiu Y, Lai Z, Bouamar H, Ramirez AG, Cigarroa FG, Sun L, Chen Y. An ancestry informative marker panel design for individual ancestry estimation of Hispanic population using whole exome sequencing data. BMC Genomics. 2019;20(S12). doi:10.1186/s12864-019-6333-6. PMID:31888480.