IntegratedPhasing

IntegratedPhasing integrates long-range Hi-C sequencing data with population-based phasing to produce more complete and accurate haplotype assemblies for human genomes.


Key Features:

  • Integration of Long-Range and Short-Range Haplotype Information: Combines long-range haplotype information from Hi-C sequencing with short-range haplotype data from population reference panels using a likelihood-based approach to improve completeness and accuracy.
  • Statistical Phasing Method: Optimizes a second-order approximation of population-based haplotype likelihood and encodes the population-based likelihood as pseudo-reads that are used alongside sequence reads.
  • Maximum Spanning Tree Algorithm: Employs a maximum spanning tree algorithm to determine the optimal haplotype configuration.
  • Compatibility with Existing Tools: Leverages HapCUT2 for the final haplotype assembly step.

Scientific Applications:

  • Hi-C Data Performance: On whole-genome Hi-C data for human genomes NA19240 and NA12878, phased 97–98% of variants and reduced switch error rates by 3–6-fold compared to existing methods.
  • Strand-seq Data Enhancement: Applied to Strand-seq data for NA12878, increased haplotype completeness from 71.4% to 94.6% and halved switch error rates.

Methodology:

Integrates Hi-C sequencing data with population reference panel haplotypes, applies a likelihood-based statistical phasing that optimizes a second-order approximation with population likelihoods encoded as pseudo-reads, uses a maximum spanning tree to select haplotype configuration, and finalizes assembly using HapCUT2.

Topics

Details

Programming Languages:
Python
Added:
11/14/2019
Last Updated:
1/9/2021

Operations

Publications

Bansal V. Integrating read-based and population-based phasing for dense and accurate haplotyping of individual genomes. Bioinformatics. 2019;35(14):i242-i248. doi:10.1093/bioinformatics/btz329. PMID:31510646. PMCID:PMC6612846.