FQSqueezer

FQSqueezer compresses FASTQ files using k-mer statistics and algorithms inspired by prediction by partial matching and dynamic Markov coder algorithms to reduce storage requirements of sequencing data generated by modern sequencing instruments.


Key Features:

  • FASTQ compression: Compresses sequencing reads stored in FASTQ format to reduce data size.
  • K-mer statistics: Leverages k-mer statistics for modeling sequence content.
  • PPM- and DMC-inspired algorithm: Employs an algorithm inspired by prediction by partial matching and dynamic Markov coder algorithms for compression.
  • Read support: Handles both single-end and paired-end reads of variable lengths.
  • High compression ratio: Achieves compression ratios often tens of percent superior to state-of-the-art tools.
  • Resource requirements: Requires considerable memory and computational time resources.

Scientific Applications:

  • High-throughput sequencing data management: Reduces the storage footprint of sequencing datasets produced by modern sequencing instruments.
  • Large-scale genomic studies: Enables more efficient storage for large-scale genomic projects.
  • Archival storage: Supports long-term archival of FASTQ data by maximizing compression.
  • Storage cost reduction: Lowers storage-related resource demands in sequencing workflows.

Methodology:

Leverages k-mer statistics and an algorithm inspired by prediction by partial matching and dynamic Markov coder algorithms.

Topics

Details

Programming Languages:
C++, C
Added:
1/18/2021
Last Updated:
3/11/2021

Operations

Publications

Deorowicz S. FQSqueezer: k-mer-based compression of sequencing data. Scientific Reports. 2020;10(1). doi:10.1038/s41598-020-57452-6. PMID:31953467. PMCID:PMC6969201.

PMID: 31953467
PMCID: PMC6969201
Funding: - Narodowe Centrum Nauki: DEC-2015/17/B/ST6/01890 - Ministry of Science and Higher Education | Narodowe Centrum Badań i Rozwoju: POIG.02.03.01-24-099/13