contiBAIT

contiBAIT refines de novo genome assemblies by integrating Strand-seq single-cell strand inheritance data to detect and correct scaffold misorientations and mislocalizations and to cluster and order contigs into putative chromosomes.


Key Features:

  • Implementation: Implemented as an R/Bioconductor package.
  • Strand Inheritance Data Utilization: Uses Strand-seq single-cell strand inheritance data to establish relative orientations and distances between scaffolds.
  • Error Correction: Identifies and corrects scaffold misorientations and mislocalizations arising from repetitive regions and sequencing errors.
  • Chromosome-Level Assembly Improvement: Clusters unbridged contigs into putative chromosomes and orders scaffolds within chromosomes.
  • Complementary Methodology: Provides orthogonal verification to conventional sequencing-based assemblies by leveraging strand inheritance signals.

Scientific Applications:

  • De novo assembly refinement: Improves orientation and placement of scaffolds in de novo genome assemblies.
  • Resolving complex regions: Helps resolve repetitive and structurally complex genomic regions that cause assembly errors in conventional sequencing.
  • Chromosome-scale builds and downstream analyses: Facilitates generation of chromosome-scale builds and reduces structural assembly errors that affect downstream genetic analyses.

Methodology:

Integrates Strand-seq single-cell strand inheritance data with existing genome assemblies to infer relative genomic distances and orientations, detect misorientations and mislocalizations, cluster contigs into putative chromosomes, and order scaffolds within those chromosomes.

Topics

Collections

Details

License:
BSD-4-Clause
Tool Type:
command-line tool, library
Operating Systems:
Linux, Windows, Mac
Programming Languages:
R
Added:
1/17/2017
Last Updated:
1/17/2019

Operations

Publications

O’Neill K, Hills M, Gottlieb M, Borkowski M, Karsan A, Lansdorp PM. Assembling draft genomes using contiBAIT. Bioinformatics. 2017;33(17):2737-2739. doi:10.1093/bioinformatics/btx281. PMID:28475666. PMCID:PMC5860061.

PMID: 28475666
PMCID: PMC5860061
Funding: - National Institutes of Health: R01GM094146 - Terry Fox Research Institute: TFF-122869

Documentation

Downloads