HARSH
HARSH infers haplotypes by integrating multi-SNP short read sequencing data with reference haplotype panels to improve phasing accuracy for genetic linkage and variation analyses.
Key Features:
- Integration of multi-SNP reads: Incorporates information from short reads that span multiple SNPs to capture read-level haplotypic origin data.
- Reference-panel conditioning: Conditions haplotype inference on a reference haplotype panel to identify explanatory haplotype segments.
- Probabilistic model: Uses a novel probabilistic model to identify the most likely haplotype segments from the reference panel that explain an individual's sequencing data.
- Gibbs sampler: Employs an efficient Gibbs sampler-based sampling method to perform inference.
- Relation to HMM/copying models: Builds on the haplotype copying/HMM framework for representing individual haplotypes as mosaics of reference haplotypes.
- Empirical performance: Demonstrates approximately 20% higher accuracy versus basic haplotype copying models and about 10% higher accuracy versus Hap-SeqX, with reduced computational time and memory.
- Benchmark datasets: Validated using simulated sequencing reads derived from real genotypes in the HapMap project and the 1000 Genomes Project.
Scientific Applications:
- Haplotype phasing from sequencing: Phasing of individual genomes using short read sequencing data combined with reference panels.
- Improved genotype calling and phasing accuracy: Enhancing accuracy of genotype calls and phased haplotypes by leveraging multi-SNP read information.
- Population genetics analyses: Facilitating analyses of genetic linkage and variation using HapMap and 1000 Genomes reference panels.
Methodology:
Implements a probabilistic model that integrates multi-SNP short read data and conditions on a reference haplotype panel to identify likely reference haplotype segments using an efficient Gibbs sampler.
Topics
Details
- Tool Type:
- command-line tool
- Operating Systems:
- Linux, Windows, Mac
- Programming Languages:
- Python
- Added:
- 12/18/2017
- Last Updated:
- 11/24/2024
Operations
Publications
Yang W, Hormozdiari F, Wang Z, He D, Pasaniuc B, Eskin E. Leveraging reads that span multiple single nucleotide polymorphisms for haplotype inference from sequencing data. Bioinformatics. 2013;29(18):2245-2252. doi:10.1093/bioinformatics/btt386. PMID:23825370. PMCID:PMC3753566.