IntegratedPhasing
IntegratedPhasing integrates long-range Hi-C sequencing data with population-based phasing to produce more complete and accurate haplotype assemblies for human genomes.
Key Features:
- Integration of Long-Range and Short-Range Haplotype Information: Combines long-range haplotype information from Hi-C sequencing with short-range haplotype data from population reference panels using a likelihood-based approach to improve completeness and accuracy.
- Statistical Phasing Method: Optimizes a second-order approximation of population-based haplotype likelihood and encodes the population-based likelihood as pseudo-reads that are used alongside sequence reads.
- Maximum Spanning Tree Algorithm: Employs a maximum spanning tree algorithm to determine the optimal haplotype configuration.
- Compatibility with Existing Tools: Leverages HapCUT2 for the final haplotype assembly step.
Scientific Applications:
- Hi-C Data Performance: On whole-genome Hi-C data for human genomes NA19240 and NA12878, phased 97–98% of variants and reduced switch error rates by 3–6-fold compared to existing methods.
- Strand-seq Data Enhancement: Applied to Strand-seq data for NA12878, increased haplotype completeness from 71.4% to 94.6% and halved switch error rates.
Methodology:
Integrates Hi-C sequencing data with population reference panel haplotypes, applies a likelihood-based statistical phasing that optimizes a second-order approximation with population likelihoods encoded as pseudo-reads, uses a maximum spanning tree to select haplotype configuration, and finalizes assembly using HapCUT2.
Topics
Details
- Programming Languages:
- Python
- Added:
- 11/14/2019
- Last Updated:
- 1/9/2021
Operations
Publications
Bansal V. Integrating read-based and population-based phasing for dense and accurate haplotyping of individual genomes. Bioinformatics. 2019;35(14):i242-i248. doi:10.1093/bioinformatics/btz329. PMID:31510646. PMCID:PMC6612846.