FastqCLS

FastqCLS compresses FASTQ files from long-read sequencing to reduce storage and transfer requirements for genomic data.


Key Features:

  • Compression algorithm: A specialized algorithm that targets FASTQ files from long-read sequencing to achieve significant file size reductions.
  • Read reordering with scoring model: Read reordering using a novel scoring model to improve compressibility while preserving all original information (lossless).
  • Long-read optimization: Optimization for the specific characteristics and demands of long-read sequencing data compared with techniques developed for short-read sequencing.
  • Integrated processing: Incorporates the necessary data processing steps into a single software package.
  • Benchmark performance: Demonstrated superior compression ratios versus major FASTQ compression tools on benchmark datasets, including newly generated long-read sequencing data.

Scientific Applications:

  • Genomic data storage: Reducing storage and transfer burdens for genome sequencing data produced by long-read sequencing technologies.
  • Long-read dataset management: Enabling more efficient management of long-read sequencing datasets in genomics research and downstream analysis workflows.
  • Method comparison: Serving as a reference approach for comparative evaluations of FASTQ compressors using benchmark long-read datasets.

Methodology:

Read reordering using a novel scoring model and application of a compression algorithm to FASTQ files; comparative evaluation against major FASTQ compression tools using benchmark datasets including newly generated long-read sequencing data.

Topics

Details

License:
MIT
Cost:
Free of charge
Tool Type:
command-line tool
Operating Systems:
Mac, Linux, Windows
Programming Languages:
Python
Added:
5/9/2022
Last Updated:
5/9/2022

Operations

Data Inputs & Outputs

Data handling

Outputs

    Publications

    Lee D, Song G. FastqCLS: a FASTQ compressor for long-read sequencing via read reordering using a novel scoring model. Bioinformatics. 2021;38(2):351-356. doi:10.1093/bioinformatics/btab696. PMID:34623374.

    PMID: 34623374
    Funding: - Korea government: 2020-0-01450 - National Research Foundation of Korea (NRF) grant funded by the Korea government: 2021R1A2C2010775