FQSqueezer
FQSqueezer compresses FASTQ files using k-mer statistics and algorithms inspired by prediction by partial matching and dynamic Markov coder algorithms to reduce storage requirements of sequencing data generated by modern sequencing instruments.
Key Features:
- FASTQ compression: Compresses sequencing reads stored in FASTQ format to reduce data size.
- K-mer statistics: Leverages k-mer statistics for modeling sequence content.
- PPM- and DMC-inspired algorithm: Employs an algorithm inspired by prediction by partial matching and dynamic Markov coder algorithms for compression.
- Read support: Handles both single-end and paired-end reads of variable lengths.
- High compression ratio: Achieves compression ratios often tens of percent superior to state-of-the-art tools.
- Resource requirements: Requires considerable memory and computational time resources.
Scientific Applications:
- High-throughput sequencing data management: Reduces the storage footprint of sequencing datasets produced by modern sequencing instruments.
- Large-scale genomic studies: Enables more efficient storage for large-scale genomic projects.
- Archival storage: Supports long-term archival of FASTQ data by maximizing compression.
- Storage cost reduction: Lowers storage-related resource demands in sequencing workflows.
Methodology:
Leverages k-mer statistics and an algorithm inspired by prediction by partial matching and dynamic Markov coder algorithms.
Topics
Details
- Programming Languages:
- C++, C
- Added:
- 1/18/2021
- Last Updated:
- 3/11/2021
Operations
Publications
Deorowicz S. FQSqueezer: k-mer-based compression of sequencing data. Scientific Reports. 2020;10(1). doi:10.1038/s41598-020-57452-6. PMID:31953467. PMCID:PMC6969201.
PMID: 31953467
PMCID: PMC6969201
Funding: - Narodowe Centrum Nauki: DEC-2015/17/B/ST6/01890
- Ministry of Science and Higher Education | Narodowe Centrum Badań i Rozwoju: POIG.02.03.01-24-099/13