ERStruct
ERStruct infers population structure from large-scale, multi-ethnic whole genome sequencing (WGS) data to enable analysis of genetic diversity and population stratification.
Key Features:
- Parallel computing and GPU acceleration: Implements optimized matrix operations using parallel processing and GPU acceleration to speed computations on large-scale datasets.
- Adaptive data splitting: Splits input data adaptively to accommodate GPUs with limited memory while enabling processing of extensive genomic datasets.
- Estimation of informative principal components: Estimates the number of top principal components that are most informative for capturing population structure.
- PCA tailored for ultra-dimensional WGS data: Adapts principal component analysis to handle ultra-dimensional whole genome sequencing data and complex linkage disequilibrium patterns.
Scientific Applications:
- Population genetics: Infers population structure and genetic diversity within and between populations from WGS data.
- Human ancestry analysis: Supports inference of ancestry and relationships among multi-ethnic populations using genomic principal components.
- Disease association mapping: Provides population structure estimates for controlling stratification in genome-wide association studies (GWAS).
- Personalized medicine: Informs interpretation of genetic variation across ethnic groups relevant to precision medicine.
Methodology:
Uses principal component analysis (PCA) adapted for ultra-dimensional genomic data, estimates the number of informative principal components, and employs optimized matrix operations via parallel processing and GPU acceleration with adaptive data splitting to accommodate limited GPU memory.
Topics
Details
- License:
- MIT
- Cost:
- Free of charge
- Tool Type:
- library
- Operating Systems:
- Mac, Linux, Windows
- Programming Languages:
- Python
- Added:
- 1/2/2024
- Last Updated:
- 11/24/2024
Operations
Publications
Yang J, Xu Y, Yao M, Wang G, Liu Z. ERStruct: a fast Python package for inferring the number of top principal components from whole genome sequencing data. BMC Bioinformatics. 2023;24(1). doi:10.1186/s12859-023-05305-0. PMID:37131141. PMCID:PMC10155328.