GenCompress

GenCompress compresses DNA sequences losslessly by identifying and exploiting approximate repeats to reduce storage requirements and support analyses of sequence relatedness.


Key Features:

  • Lossless compression: Preserves all nucleotide information while reducing DNA sequence size.
  • Approximate repeat detection: Systematically searches for patterns that recur with slight variations across sequences.
  • Redundancy exploitation: Exploits subtle sequence redundancies to achieve improved compression.
  • Benchmark performance: Demonstrates superior compression ratios on benchmark DNA sequences compared with existing methods.
  • Sequence relatedness metric: Implements a theoretical framework that measures relatedness between two DNA sequences.
  • Experimental validation: Algorithm behavior and the theoretical framework have been validated through experiments on DNA sequence datasets.

Scientific Applications:

  • Comparative genomics: Enables efficient storage and comparison of genomes across species or within populations.
  • Phylogenetics: Assists in constructing evolutionary trees and analyzing evolutionary relationships.
  • Genomic data management: Reduces storage and retrieval burden for large genomic datasets.

Methodology:

Systematic search for approximate repeats and exploitation of these subtle redundancies to perform lossless compression; application of a theoretical framework to measure relatedness between two DNA sequences; experimental validation on benchmark DNA sequences.

Topics

Details

Tool Type:
command-line tool
Operating Systems:
Linux, Windows
Programming Languages:
Perl
Added:
8/3/2017
Last Updated:
12/10/2018

Operations

Publications

Chen X, et al. A Compression Algorithm for DNA Sequences and Its Applications in Genome Comparison. Genome Inform Ser Workshop Genome Inform. 1999; 10:51-61.

PMID: 11072342

Documentation

Links