PMFFRC

PMFFRC optimizes compression of genomic short reads by leveraging memory modeling and redundant clustering to reduce storage requirements for high-throughput sequencing datasets.


Key Features:

  • Optimization Objective: PMFFRC focuses on maximizing compression ratio for genomic short reads to reduce storage space for sequencing data.
  • Memory Modeling and Redundant Clustering: The tool employs memory modeling for big-memory systems and redundant clustering to identify and exploit duplicative information across sequencing files.
  • Performance Gains: Applied to 982 GB of fastq format sequencing data containing 274 GB and 3.3 billion short reads, PMFFRC achieved average maximum compression ratio gains of 77.89%, 77.56%, 73.51%, and 29.36% over HARC, SPRING, Mstcom, and FastqCLS, respectively.
  • Storage Space Savings: PMFFRC reduced storage space requirements by 39.41%, 41.62%, 40.99%, and 20.19% compared to HARC, SPRING, Mstcom, and FastqCLS, respectively.
  • Resource Efficiency: The tool rationalizes memory usage on compression servers to improve resource efficiency.

Scientific Applications:

  • High-throughput sequencing data storage and sharing: Enables reduced storage requirements and more efficient sharing of large-scale sequencing datasets in environments with routine data generation.
  • Genomics, personalized medicine, and bioinformatics: Supports research in genomics, personalized medicine, and bioinformatics by lowering storage barriers and facilitating broader data accessibility and collaboration.

Methodology:

Memory modeling to optimize use of large memory systems; redundant clustering to identify and cluster redundant information across sequencing files.

Topics

Details

License:
Apache-2.0
Cost:
Free of charge
Tool Type:
command-line tool
Operating Systems:
Linux
Programming Languages:
Shell, C++, Python
Added:
4/19/2024
Last Updated:
11/24/2024

Operations

Publications

Sun H, Zheng Y, Xie H, Ma H, Liu X, Wang G. PMFFRC: a large-scale genomic short reads compression optimizer via memory modeling and redundant clustering. BMC Bioinformatics. 2023;24(1). doi:10.1186/s12859-023-05566-9. PMID:38036969. PMCID:PMC10691058.

Downloads