HAPNEST

HAPNEST generates synthetic genotype and phenotype datasets for large-scale genetic analyses by producing individual-level genotypic and phenotypic data that preserve key statistical properties while controlling relatedness to reference panels.


Key Features:

  • Scalability and Efficiency: Capable of producing synthetic datasets for up to 1 million individuals with faster computational performance compared to alternative methods.
  • Statistical Fidelity: Preserves essential statistical properties of real genotype and phenotype data to support valid research analyses.
  • Low Relatedness: Generates synthetic datasets with a lower degree of relatedness to existing reference panels to enhance dataset diversity.
  • Versatility in Data Generation: Can simulate 6.8 million common variants and nine phenotypes with varying degrees of heritability and polygenicity.

Scientific Applications:

  • Biobank-scale analyses: Facilitates comparison of methods for generating polygenic risk scores across multiple ancestry groups and varied genetic architectures.
  • Benchmarking and method development: Provides synthetic data for evaluating synthetic data quality and benchmarking polygenic risk score methodologies for complex traits and diseases.

Methodology:

Implemented in Julia and C.

Topics

Details

License:
CC-BY-4.0
Cost:
Free of charge
Tool Type:
workflow
Programming Languages:
Julia, C
Added:
3/6/2024
Last Updated:
11/24/2024

Operations

Publications

Wharrie S, Yang Z, Raj V, Monti R, Gupta R, Wang Y, Martin A, O’Connor LJ, Kaski S, Marttinen P, Palamara PF, Lippert C, Ganna A. HAPNEST: efficient, large-scale generation and evaluation of synthetic datasets for genotypes and phenotypes. Bioinformatics. 2023;39(9). doi:10.1093/bioinformatics/btad535. PMID:37647640. PMCID:PMC10493177.

PMID: 37647640
Funding: - European Union’s Horizon 2020 research and innovation programme: 101016775