ERStruct

ERStruct infers population structure from large-scale, multi-ethnic whole genome sequencing (WGS) data to enable analysis of genetic diversity and population stratification.


Key Features:

  • Parallel computing and GPU acceleration: Implements optimized matrix operations using parallel processing and GPU acceleration to speed computations on large-scale datasets.
  • Adaptive data splitting: Splits input data adaptively to accommodate GPUs with limited memory while enabling processing of extensive genomic datasets.
  • Estimation of informative principal components: Estimates the number of top principal components that are most informative for capturing population structure.
  • PCA tailored for ultra-dimensional WGS data: Adapts principal component analysis to handle ultra-dimensional whole genome sequencing data and complex linkage disequilibrium patterns.

Scientific Applications:

  • Population genetics: Infers population structure and genetic diversity within and between populations from WGS data.
  • Human ancestry analysis: Supports inference of ancestry and relationships among multi-ethnic populations using genomic principal components.
  • Disease association mapping: Provides population structure estimates for controlling stratification in genome-wide association studies (GWAS).
  • Personalized medicine: Informs interpretation of genetic variation across ethnic groups relevant to precision medicine.

Methodology:

Uses principal component analysis (PCA) adapted for ultra-dimensional genomic data, estimates the number of informative principal components, and employs optimized matrix operations via parallel processing and GPU acceleration with adaptive data splitting to accommodate limited GPU memory.

Topics

Details

License:
MIT
Cost:
Free of charge
Tool Type:
library
Operating Systems:
Mac, Linux, Windows
Programming Languages:
Python
Added:
1/2/2024
Last Updated:
11/24/2024

Operations

Publications

Yang J, Xu Y, Yao M, Wang G, Liu Z. ERStruct: a fast Python package for inferring the number of top principal components from whole genome sequencing data. BMC Bioinformatics. 2023;24(1). doi:10.1186/s12859-023-05305-0. PMID:37131141. PMCID:PMC10155328.

PMID: 37131141
Funding: - National Institutes of Health: R01 AG076901