ORNA

ORNA performs read normalization of high-coverage sequencing datasets to reduce redundancy while preserving k-mers and their relative abundances to maintain de Bruijn graph connectivity for genome and transcriptome assembly.


Key Features:

  • Set Multi-Cover Optimization: Formulates normalization as a set multi-cover problem that minimizes the number of reads required while ensuring retention of original k-mers and their relative abundances and preserving de Bruijn graph connections.
  • Preservation of K-Mer Abundances: Guarantees retention of all k-mers and their relative abundances from the original dataset.
  • Performance and Efficiency: Reduces dataset size and computational resource requirements with reported assembly quality loss between 1% and 30% depending on reduction stringency.
  • Error-Corrected Data Application: Normalizing error-corrected reads yields more accurate assemblies than normalizing uncorrected datasets.
  • Novel Transcript Prediction: Combining and normalizing multiple RNA-seq datasets enables prediction of novel transcripts that may be missed otherwise.
  • General-Purpose Normalization: Applicable across diverse RNA-seq datasets with varying coverage to reduce redundancy while maintaining assembly-relevant information.

Scientific Applications:

  • Genome assembly: Supports de Bruijn graph-based genome assembly by preserving k-mers and graph connectivity during read reduction.
  • Transcriptome assembly: Supports transcript reconstruction by maintaining k-mer abundances critical for assembling transcripts from RNA-seq data.
  • Novel transcript discovery: Enables detection of previously unidentified transcripts by integrating and normalizing multiple datasets.
  • High-coverage data management: Reduces redundancy and computational resource demands in large-scale sequencing projects.
  • Preprocessing for error-corrected data: Enhances assembly accuracy when applied to error-corrected sequencing reads.

Methodology:

ORNA formulates read normalization as a set multi-cover optimization that selects a minimal subset of reads while guaranteeing retention of all k-mers and their relative abundances to preserve de Bruijn graph connectivity.

Topics

Details

Tool Type:
command-line tool
Operating Systems:
Linux, Mac
Programming Languages:
R
Added:
6/2/2018
Last Updated:
11/25/2024

Operations

Publications

Durai DA, Schulz MH. <i>In silico</i> read normalization using set multi-cover optimization. Bioinformatics. 2018;34(19):3273-3280. doi:10.1093/bioinformatics/bty307. PMID:29912280. PMCID:PMC6157080.

PMID: 29912280
PMCID: PMC6157080
Funding: - Cluster of Excellence on Multi-modal Computing and Interaction: EXC284

Documentation