CURC
CURC compresses high-throughput sequencing FASTQ files containing nucleotide sequences and quality scores using a CUDA-accelerated, reference-free, lossless pseudogenome approach to improve storage and transfer efficiency.
Key Features:
- CUDA acceleration: Uses the CUDA library to perform compression computations on GPUs.
- Heterogeneous parallel scheme: Combines GPU and CPU resources for parallelized processing.
- Reference-free compression: Operates without requiring an external reference genome.
- Lossless pseudogenome approach: Implements a pseudogenome-based method to ensure no data loss during compression.
- FASTQ-focused: Targets FASTQ files that store nucleotide sequences and quality scores from high-throughput sequencing (HTS).
- Speedup versus CPU-based compressors: Reports a 2–6-fold acceleration in compression times compared to traditional CPU-only compressors.
- Competitive compression rates: Achieves compression ratios comparable to state-of-the-art reference-free read compressors.
Scientific Applications:
- HTS data storage and transfer: Reduces storage footprint and bandwidth requirements for high-throughput sequencing datasets.
- Large-scale sequencing projects: Enables more efficient handling and archiving of large sequencing runs and cohort datasets.
- Preservation of sequencing information: Maintains nucleotide sequences and quality scores for downstream genomic analyses through lossless compression.
Methodology:
CURC implements a reference-free, lossless pseudogenome-based compression of FASTQ reads using the CUDA library for GPU acceleration within a heterogeneous parallel scheme that combines GPU and CPU resources.
Topics
Details
- License:
- MIT
- Cost:
- Free of charge
- Tool Type:
- command-line tool
- Operating Systems:
- Linux
- Programming Languages:
- C++
- Added:
- 8/15/2022
- Last Updated:
- 11/24/2024
Operations
Publications
Xie S, He X, He S, Zhu Z. CURC: a CUDA-based reference-free read compressor. Bioinformatics. 2022;38(12):3294-3296. doi:10.1093/bioinformatics/btac333. PMID:35579371.
PMID: 35579371
Funding: - National Key Research and Development Project: 2019YFE0109600
- National Natural Science Foundation of China: 61871272
- Shenzhen Fundamental Research Program: JCYJ20190808173617147
- BGIShenzhen: BGIRSZ20200002