HAPNEST
HAPNEST generates synthetic genotype and phenotype datasets for large-scale genetic analyses by producing individual-level genotypic and phenotypic data that preserve key statistical properties while controlling relatedness to reference panels.
Key Features:
- Scalability and Efficiency: Capable of producing synthetic datasets for up to 1 million individuals with faster computational performance compared to alternative methods.
- Statistical Fidelity: Preserves essential statistical properties of real genotype and phenotype data to support valid research analyses.
- Low Relatedness: Generates synthetic datasets with a lower degree of relatedness to existing reference panels to enhance dataset diversity.
- Versatility in Data Generation: Can simulate 6.8 million common variants and nine phenotypes with varying degrees of heritability and polygenicity.
Scientific Applications:
- Biobank-scale analyses: Facilitates comparison of methods for generating polygenic risk scores across multiple ancestry groups and varied genetic architectures.
- Benchmarking and method development: Provides synthetic data for evaluating synthetic data quality and benchmarking polygenic risk score methodologies for complex traits and diseases.
Methodology:
Implemented in Julia and C.
Topics
Details
- License:
- CC-BY-4.0
- Cost:
- Free of charge
- Tool Type:
- workflow
- Programming Languages:
- Julia, C
- Added:
- 3/6/2024
- Last Updated:
- 11/24/2024
Operations
Publications
Wharrie S, Yang Z, Raj V, Monti R, Gupta R, Wang Y, Martin A, O’Connor LJ, Kaski S, Marttinen P, Palamara PF, Lippert C, Ganna A. HAPNEST: efficient, large-scale generation and evaluation of synthetic datasets for genotypes and phenotypes. Bioinformatics. 2023;39(9). doi:10.1093/bioinformatics/btad535. PMID:37647640. PMCID:PMC10493177.
PMID: 37647640
PMCID: PMC10493177
Funding: - European Union’s Horizon 2020 research and innovation programme: 101016775