KOMPUTE
KOMPUTE imputes missing phenotype summary statistics in high-throughput model organism gene-knockout datasets to enhance functional characterization of mammalian protein-coding genes.
Key Features:
- Imputation Methodology: Employs conditional distribution properties of multivariate normal distributions to estimate association Z-scores for unmeasured phenotypes by computing their conditional expectation given observed Z-scores.
- Performance versus Matrix Completion: Demonstrates superior performance compared to singular value decomposition matrix completion in evaluations on simulated and real-world datasets.
Scientific Applications:
- Gene-Phenotype Association Studies: Enables more complete association analyses between loss-of-function genotypes from IMPC gene-knockout studies and phenotypic outcomes by imputing missing summary statistics.
- Enhancing Data Completeness: Fills gaps in datasets with widespread missingness, for example addressing ~75.6% missing association summary statistics reported in IMPC release version 16, to increase coverage for downstream analyses.
Methodology:
KOMPUTE uses the conditional expectation of multivariate normal distributions to predict missing association Z-scores based on observed phenotype Z-scores.
Topics
Details
- Cost:
- Free of charge
- Tool Type:
- library
- Programming Languages:
- R
- Added:
- 11/7/2023
- Last Updated:
- 11/24/2024
Operations
Publications
Warkentin C, O’Connell MJ, Lee D. KOMPUTE: imputing summary statistics of missing phenotypes in high-throughput model organism data. Bioinformatics Advances. 2023;3(1). doi:10.1093/bioadv/vbad100. PMID:37565237. PMCID:PMC10409646.
Documentation
Training material
https://statsleelab.github.io/komputeExamples