HiC-Hiker

HiC-Hiker infers contig orientations within chromosome-length scaffolds from Hi-C contact data using probabilistic modeling to improve genome assembly accuracy.


Key Features:

  • Probabilistic Modeling: Models Hi-C contact frequencies across contigs using a probabilistic framework to infer contig orientations.
  • Viterbi Algorithm: Employs the Viterbi algorithm to determine the most probable orientation sequence of contigs, reducing errors for shorter contigs with fewer neighboring contacts.
  • Error Reduction: Demonstrates reduced contig orientation error rates compared to 3D-DNA, reported from 4.3% down to 1.7%.
  • Long-Range Information Utilization: Incorporates long-range Hi-C contacts between distal contigs to refine contact probability estimates and improve contig ordering.

Scientific Applications:

  • De novo Genome Assembly: Orients contigs to produce chromosome-length scaffolds in de novo assemblies, including non-model species lacking established genome markers.
  • Genomic Distance Estimation: Uses Hi-C contact frequencies to estimate genomic distances between loci for scaffold construction.
  • Gene Prediction: Supports more accurate gene prediction by improving scaffold orientation.
  • Read Alignment: Improves read alignment accuracy by providing more accurate scaffold structures.
  • Synteny Block Estimation: Aids comparative genomics through more reliable estimation of synteny blocks.

Methodology:

Models Hi-C contact frequencies across contigs using a probabilistic framework and infers the most probable contig orientations with the Viterbi algorithm, optionally incorporating long-range Hi-C contacts to refine contact probability estimates.

Topics

Details

License:
BSD-3-Clause
Tool Type:
command-line tool
Programming Languages:
Python
Added:
1/18/2021
Last Updated:
1/30/2021

Operations

Publications

Nakabayashi R, Morishita S. HiC-Hiker: a probabilistic model to determine contig orientation in chromosome-length scaffolds with Hi-C. Bioinformatics. 2020;36(13):3966-3974. doi:10.1093/bioinformatics/btaa288. PMID:32369554. PMCID:PMC7672694.