PMFFRC
PMFFRC optimizes compression of genomic short reads by leveraging memory modeling and redundant clustering to reduce storage requirements for high-throughput sequencing datasets.
Key Features:
- Optimization Objective: PMFFRC focuses on maximizing compression ratio for genomic short reads to reduce storage space for sequencing data.
- Memory Modeling and Redundant Clustering: The tool employs memory modeling for big-memory systems and redundant clustering to identify and exploit duplicative information across sequencing files.
- Performance Gains: Applied to 982 GB of fastq format sequencing data containing 274 GB and 3.3 billion short reads, PMFFRC achieved average maximum compression ratio gains of 77.89%, 77.56%, 73.51%, and 29.36% over HARC, SPRING, Mstcom, and FastqCLS, respectively.
- Storage Space Savings: PMFFRC reduced storage space requirements by 39.41%, 41.62%, 40.99%, and 20.19% compared to HARC, SPRING, Mstcom, and FastqCLS, respectively.
- Resource Efficiency: The tool rationalizes memory usage on compression servers to improve resource efficiency.
Scientific Applications:
- High-throughput sequencing data storage and sharing: Enables reduced storage requirements and more efficient sharing of large-scale sequencing datasets in environments with routine data generation.
- Genomics, personalized medicine, and bioinformatics: Supports research in genomics, personalized medicine, and bioinformatics by lowering storage barriers and facilitating broader data accessibility and collaboration.
Methodology:
Memory modeling to optimize use of large memory systems; redundant clustering to identify and cluster redundant information across sequencing files.
Topics
Details
- License:
- Apache-2.0
- Cost:
- Free of charge
- Tool Type:
- command-line tool
- Operating Systems:
- Linux
- Programming Languages:
- Shell, C++, Python
- Added:
- 4/19/2024
- Last Updated:
- 11/24/2024
Operations
Publications
Sun H, Zheng Y, Xie H, Ma H, Liu X, Wang G. PMFFRC: a large-scale genomic short reads compression optimizer via memory modeling and redundant clustering. BMC Bioinformatics. 2023;24(1). doi:10.1186/s12859-023-05566-9. PMID:38036969. PMCID:PMC10691058.