permGWAS

permGWAS implements permutation-based significance thresholds for genome-wide association studies to improve statistical rigor in detecting genotype-phenotype associations.


Key Features:

  • Python implementation: Implemented in Python for genome-wide association analysis.
  • LMM reformulation with 4D tensors: Reformulates linear mixed models (LMM) using 4D tensors to enable efficient computation of permutation-based significance thresholds.
  • Multi-CPU and GPU acceleration: Supports multi-CPU and GPU execution to reduce computational time for large-scale permutation testing.
  • Demonstrated scalability: Demonstrated re-analysis of over 500 Arabidopsis thaliana phenotypes with 100 permutations each in under eight days on a single GPU.
  • Improved statistical control: Uses permutation-based thresholds to provide more realistic significance levels and lower false discovery rates compared to Bonferroni correction, particularly for skewed phenotypic distributions.

Scientific Applications:

  • Plant genetics (Arabidopsis thaliana): Applicable to large-scale GWAS in plant genetics, including analyses of Arabidopsis thaliana phenotypes.
  • Permutation-based GWAS significance estimation: Provides empirical significance thresholds for GWAS where LMM assumptions such as normal residuals and marker independence may be violated.
  • Large-scale phenotype analyses: Suited for extensive multi-phenotype studies requiring many permutations to control false discoveries.

Methodology:

permGWAS reformulates linear mixed models using 4D tensors to perform permutation testing and to compute permutation-based significance thresholds, with support for multi-CPU and GPU execution.

Topics

Details

License:
MIT
Cost:
Free of charge
Tool Type:
command-line tool
Programming Languages:
Python
Added:
11/6/2022
Last Updated:
11/24/2024

Operations

Publications

John M, Ankenbrand MJ, Artmann C, Freudenthal JA, Korte A, Grimm DG. Efficient permutation-based genome-wide association studies for normal and skewed phenotypic distributions. Bioinformatics. 2022;38(Supplement_2):ii5-ii12. doi:10.1093/bioinformatics/btac455. PMID:36124808. PMCID:PMC9486594.

PMID: 36124808
PMCID: PMC9486594
Funding: - Federal Ministry of Education and Research: ECCB2022 - BMBF: 01—S21038B