ORNA
ORNA performs read normalization of high-coverage sequencing datasets to reduce redundancy while preserving k-mers and their relative abundances to maintain de Bruijn graph connectivity for genome and transcriptome assembly.
Key Features:
- Set Multi-Cover Optimization: Formulates normalization as a set multi-cover problem that minimizes the number of reads required while ensuring retention of original k-mers and their relative abundances and preserving de Bruijn graph connections.
- Preservation of K-Mer Abundances: Guarantees retention of all k-mers and their relative abundances from the original dataset.
- Performance and Efficiency: Reduces dataset size and computational resource requirements with reported assembly quality loss between 1% and 30% depending on reduction stringency.
- Error-Corrected Data Application: Normalizing error-corrected reads yields more accurate assemblies than normalizing uncorrected datasets.
- Novel Transcript Prediction: Combining and normalizing multiple RNA-seq datasets enables prediction of novel transcripts that may be missed otherwise.
- General-Purpose Normalization: Applicable across diverse RNA-seq datasets with varying coverage to reduce redundancy while maintaining assembly-relevant information.
Scientific Applications:
- Genome assembly: Supports de Bruijn graph-based genome assembly by preserving k-mers and graph connectivity during read reduction.
- Transcriptome assembly: Supports transcript reconstruction by maintaining k-mer abundances critical for assembling transcripts from RNA-seq data.
- Novel transcript discovery: Enables detection of previously unidentified transcripts by integrating and normalizing multiple datasets.
- High-coverage data management: Reduces redundancy and computational resource demands in large-scale sequencing projects.
- Preprocessing for error-corrected data: Enhances assembly accuracy when applied to error-corrected sequencing reads.
Methodology:
ORNA formulates read normalization as a set multi-cover optimization that selects a minimal subset of reads while guaranteeing retention of all k-mers and their relative abundances to preserve de Bruijn graph connectivity.
Topics
Details
- Tool Type:
- command-line tool
- Operating Systems:
- Linux, Mac
- Programming Languages:
- R
- Added:
- 6/2/2018
- Last Updated:
- 11/25/2024
Operations
Publications
Durai DA, Schulz MH. <i>In silico</i> read normalization using set multi-cover optimization. Bioinformatics. 2018;34(19):3273-3280. doi:10.1093/bioinformatics/bty307. PMID:29912280. PMCID:PMC6157080.
PMID: 29912280
PMCID: PMC6157080
Funding: - Cluster of Excellence on Multi-modal Computing and Interaction: EXC284