PIP-SNP

PIP-SNP performs linkage disequilibrium-based SNP binning, k-nearest neighbors (kNN) genotype imputation, and synthetic marker synthesis to reduce genotype data dimensionality while preserving genetic information for GWAS and other genetic analyses.


Key Features:

  • Linkage Disequilibrium Bin Mapping: Uses linkage disequilibrium (LD) and a stochastic process to model SNP signals and compute autocorrelation coefficients to construct LD bins that group genetically linked SNPs for dimension reduction.
  • Genotype Imputation: Imputes missing genotype data using the k-nearest neighbors (kNN) algorithm based on similarity between SNP profiles.
  • Marker Synthesizing: Identifies optimal synthetic markers that represent each LD bin as dimension-reduced proxies of original genome-wide markers.
  • Information Conservation Evaluation: Assesses conservation of genetic information when replacing original SNP markers with synthetic markers to ensure retention of critical signals.

Scientific Applications:

  • Genome-wide association studies (GWAS): Facilitates GWAS by reducing SNP dimensionality while preserving association signals for phenotype-genotype analysis.
  • Population-scale genotype analyses: Enables analysis of large and diverse datasets, with demonstrated applicability to rice recombinant inbred lines (RIL) and HapMap project data.

Methodology:

Models SNP signals with a stochastic process and computes autocorrelation coefficients to form LD bins; performs k-nearest neighbors (kNN) imputation for missing genotypes; derives synthetic markers representing LD bins; and evaluates information conservation between original and synthetic markers.

Topics

Details

License:
GPL-3.0
Cost:
Free of charge
Tool Type:
command-line tool
Operating Systems:
Linux
Programming Languages:
C++
Added:
11/28/2021
Last Updated:
11/28/2021

Operations

Publications

Zhang W, Kang Y, Dai X, Xu S, Zhao PX. PIP-SNP: a pipeline for processing SNP data featured as linkage disequilibrium bin mapping, genotype imputing and marker synthesizing. NAR Genomics and Bioinformatics. 2021;3(3). doi:10.1093/nargab/lqab060. PMID:34235432. PMCID:PMC8256826.

PMID: 34235432
PMCID: PMC8256826
Funding: - National Science Foundation: DBI-1458130, DBI-1458597

Documentation