PgRC

PgRC compresses DNA sequencing reads by constructing a pseudogenome—an approximation of the shortest common superstring from high-quality reads—to reduce redundancy and improve storage efficiency for FASTQ datasets.


Key Features:

  • In-Memory Algorithm: PgRC processes DNA sequence data directly within system memory to perform compression and decompression operations.
  • Pseudogenome-Based Compression: PgRC constructs a pseudogenome from high-quality reads that approximates the shortest common superstring to consolidate overlaps and reduce redundancy.
  • Performance and Efficiency: Experimental evaluations report PgRC achieves up to ~15% and ~20% better compression ratio compared to SPRING and Minicom, respectively, while maintaining decompression speed comparable to those tools.
  • FASTQ Storage Reduction: PgRC significantly reduces storage requirements for FASTQ datasets without compromising decompression speed.

Scientific Applications:

  • Genomic Research: Facilitates storage and management of large-scale genomic sequencing datasets.
  • Transcriptomics: Enables efficient storage of RNA-seq datasets used for gene expression analysis.
  • Metagenomics: Supports compression of complex metagenomic sequencing datasets containing diverse microbial sequences.

Methodology:

PgRC constructs a pseudogenome from high-quality reads by approximating the shortest common superstring and performs in-memory compression and decompression of FASTQ DNA reads.

Topics

Details

Programming Languages:
C++, C
Added:
1/18/2021
Last Updated:
1/23/2021

Operations

Publications

Kowalski TM, Grabowski S. PgRC: pseudogenome-based read compressor. Bioinformatics. 2019;36(7):2082-2089. doi:10.1093/bioinformatics/btz919. PMID:31893286.

PMID: 31893286
Funding: - Polish National Centre for Research and Development: POIR.04.01.02-00-0089/17-00