vSampler

vSampler generates matched control genetic variants for enrichment and calibration analyses by sampling controls matched on genomic properties such as minor allele frequency (MAF), distance to the closest transcription start site (DTCT), linkage disequilibrium (LD) patterns, GC content, and tissue/cell type-specific epigenomic features.


Key Features:

  • Comprehensive Matching Properties: vSampler randomly samples control variants matched to input variants on properties including MAF, distance to the closest transcription start site (DTCT), number of nearby genes, number of variants in linkage disequilibrium (LD), GC content, and tissue/cell type-specific epigenomic features.
  • Support for SNPs and Indels: The tool handles both single nucleotide polymorphisms (SNPs) and insertions/deletions (indels).
  • Advanced Variant Annotations: vSampler provides variant annotations to accompany sampled controls for downstream analyses.
  • Performance Efficiency: Sampling is implemented using a novel data structure and sampling algorithms to improve speed on large genomic datasets.

Scientific Applications:

  • Enrichment Analysis Calibration: vSampler constructs matched-control null distributions to empirically estimate significance in GWAS and QTL enrichment analyses, mitigating confounding by variant properties.
  • Negative Data Construction: vSampler generates negative training and testing datasets for regulatory variant prediction methods by sampling non-functional matched controls.

Methodology:

Sampling uses a novel data structure and sampling algorithms to select control variants matched on specified properties (MAF, DTCT, nearby gene count, LD variant count, GC content, tissue/cell type-specific epigenomic features); it supports SNPs and indels and produces variant annotations.

Topics

Details

Added:
1/18/2021
Last Updated:
3/13/2021

Operations

Publications

Huang D, Wang Z, Zhou Y, Liang Q, Sham PC, Yao H, Li MJ. vSampler: fast and annotation-based matched variant sampling tool. Bioinformatics. 2020;37(13):1915-1917. doi:10.1093/bioinformatics/btaa883. PMID:33270826.

PMID: 33270826
Funding: - National Natural Science Foundation of China: 31871327, 32070675