PIP-SNP
PIP-SNP performs linkage disequilibrium-based SNP binning, k-nearest neighbors (kNN) genotype imputation, and synthetic marker synthesis to reduce genotype data dimensionality while preserving genetic information for GWAS and other genetic analyses.
Key Features:
- Linkage Disequilibrium Bin Mapping: Uses linkage disequilibrium (LD) and a stochastic process to model SNP signals and compute autocorrelation coefficients to construct LD bins that group genetically linked SNPs for dimension reduction.
- Genotype Imputation: Imputes missing genotype data using the k-nearest neighbors (kNN) algorithm based on similarity between SNP profiles.
- Marker Synthesizing: Identifies optimal synthetic markers that represent each LD bin as dimension-reduced proxies of original genome-wide markers.
- Information Conservation Evaluation: Assesses conservation of genetic information when replacing original SNP markers with synthetic markers to ensure retention of critical signals.
Scientific Applications:
- Genome-wide association studies (GWAS): Facilitates GWAS by reducing SNP dimensionality while preserving association signals for phenotype-genotype analysis.
- Population-scale genotype analyses: Enables analysis of large and diverse datasets, with demonstrated applicability to rice recombinant inbred lines (RIL) and HapMap project data.
Methodology:
Models SNP signals with a stochastic process and computes autocorrelation coefficients to form LD bins; performs k-nearest neighbors (kNN) imputation for missing genotypes; derives synthetic markers representing LD bins; and evaluates information conservation between original and synthetic markers.
Topics
Details
- License:
- GPL-3.0
- Cost:
- Free of charge
- Tool Type:
- command-line tool
- Operating Systems:
- Linux
- Programming Languages:
- C++
- Added:
- 11/28/2021
- Last Updated:
- 11/28/2021
Operations
Publications
Zhang W, Kang Y, Dai X, Xu S, Zhao PX. PIP-SNP: a pipeline for processing SNP data featured as linkage disequilibrium bin mapping, genotype imputing and marker synthesizing. NAR Genomics and Bioinformatics. 2021;3(3). doi:10.1093/nargab/lqab060. PMID:34235432. PMCID:PMC8256826.