CURC

CURC compresses high-throughput sequencing FASTQ files containing nucleotide sequences and quality scores using a CUDA-accelerated, reference-free, lossless pseudogenome approach to improve storage and transfer efficiency.


Key Features:

  • CUDA acceleration: Uses the CUDA library to perform compression computations on GPUs.
  • Heterogeneous parallel scheme: Combines GPU and CPU resources for parallelized processing.
  • Reference-free compression: Operates without requiring an external reference genome.
  • Lossless pseudogenome approach: Implements a pseudogenome-based method to ensure no data loss during compression.
  • FASTQ-focused: Targets FASTQ files that store nucleotide sequences and quality scores from high-throughput sequencing (HTS).
  • Speedup versus CPU-based compressors: Reports a 2–6-fold acceleration in compression times compared to traditional CPU-only compressors.
  • Competitive compression rates: Achieves compression ratios comparable to state-of-the-art reference-free read compressors.

Scientific Applications:

  • HTS data storage and transfer: Reduces storage footprint and bandwidth requirements for high-throughput sequencing datasets.
  • Large-scale sequencing projects: Enables more efficient handling and archiving of large sequencing runs and cohort datasets.
  • Preservation of sequencing information: Maintains nucleotide sequences and quality scores for downstream genomic analyses through lossless compression.

Methodology:

CURC implements a reference-free, lossless pseudogenome-based compression of FASTQ reads using the CUDA library for GPU acceleration within a heterogeneous parallel scheme that combines GPU and CPU resources.

Topics

Details

License:
MIT
Cost:
Free of charge
Tool Type:
command-line tool
Operating Systems:
Linux
Programming Languages:
C++
Added:
8/15/2022
Last Updated:
11/24/2024

Operations

Publications

Xie S, He X, He S, Zhu Z. CURC: a CUDA-based reference-free read compressor. Bioinformatics. 2022;38(12):3294-3296. doi:10.1093/bioinformatics/btac333. PMID:35579371.

PMID: 35579371
Funding: - National Key Research and Development Project: 2019YFE0109600 - National Natural Science Foundation of China: 61871272 - Shenzhen Fundamental Research Program: JCYJ20190808173617147 - BGIShenzhen: BGIRSZ20200002