PgRC
PgRC compresses DNA sequencing reads by constructing a pseudogenome—an approximation of the shortest common superstring from high-quality reads—to reduce redundancy and improve storage efficiency for FASTQ datasets.
Key Features:
- In-Memory Algorithm: PgRC processes DNA sequence data directly within system memory to perform compression and decompression operations.
- Pseudogenome-Based Compression: PgRC constructs a pseudogenome from high-quality reads that approximates the shortest common superstring to consolidate overlaps and reduce redundancy.
- Performance and Efficiency: Experimental evaluations report PgRC achieves up to ~15% and ~20% better compression ratio compared to SPRING and Minicom, respectively, while maintaining decompression speed comparable to those tools.
- FASTQ Storage Reduction: PgRC significantly reduces storage requirements for FASTQ datasets without compromising decompression speed.
Scientific Applications:
- Genomic Research: Facilitates storage and management of large-scale genomic sequencing datasets.
- Transcriptomics: Enables efficient storage of RNA-seq datasets used for gene expression analysis.
- Metagenomics: Supports compression of complex metagenomic sequencing datasets containing diverse microbial sequences.
Methodology:
PgRC constructs a pseudogenome from high-quality reads by approximating the shortest common superstring and performs in-memory compression and decompression of FASTQ DNA reads.
Topics
Details
- Programming Languages:
- C++, C
- Added:
- 1/18/2021
- Last Updated:
- 1/23/2021
Operations
Publications
Kowalski TM, Grabowski S. PgRC: pseudogenome-based read compressor. Bioinformatics. 2019;36(7):2082-2089. doi:10.1093/bioinformatics/btz919. PMID:31893286.
PMID: 31893286
Funding: - Polish National Centre for Research and Development: POIR.04.01.02-00-0089/17-00