UnderIICRMS

UnderIICRMS performs alignment-free comparisons of cis-regulatory modules (CRMs) and unassembled genomes using variable-length patterns to detect regulatory sequence similarity and compare NGS read datasets.


Key Features:

  • Parameter-Free Approach: Operates without predefined parameters, removing the need for parameter tuning across datasets and biological contexts.
  • Variable-Length Patterns: Utilizes variable-length words/patterns instead of fixed-length k-mers to capture length variability in CRMs and genomic sequences.
  • Enhancer Clustering and Functional Correlation: Identifies functionally similar enhancers and clusters within cis-regulatory modules by employing a similarity measure that detects related regulatory sequences.
  • Discrimination of Functionally Related Enhancers: Distinguishes enhancers active in different tissues, as demonstrated on simulated data and ChIP-seq sequences including mouse enhancers.
  • Genome Comparison Without Assembly: Enables direct comparison of unassembled genomes using NGS reads when reference genomes are unavailable.
  • Similarity Measure Incorporating Reverses and Reverse-Complements: Defines a similarity measure based on variable-length patterns including reverses and reverse-complements and their statistical and syntactical properties.
  • Performance in Genome Discrimination: Evaluated against other alignment-free statistics on simulated and real genomes and shown to outperform fixed-length pattern methods, particularly on real genome NGS read data.

Scientific Applications:

  • Comparative Genomic Studies: Facilitates assembly-free comparison of genomes from NGS reads, applicable to non-model organisms and newly sequenced species.
  • Regulatory Sequence Analysis: Supports analysis of enhancers and CRM architecture to infer functional correlations among regulatory elements.
  • Evolutionary Biology: Enables study of evolutionary relationships between genomes using alignment-free comparisons of NGS data in the absence of a reference genome.

Methodology:

Performs alignment-free sequence comparisons using variable-length patterns/words in a parameter-free framework; computes a similarity measure that includes reverses and reverse-complements and assesses statistical and syntactical properties; validated on simulated data and ChIP-seq sequences and compared against other alignment-free statistics using unassembled NGS reads.

Topics

Details

License:
Other
Tool Type:
command-line tool
Operating Systems:
Linux, Windows, Mac
Programming Languages:
Java
Added:
1/22/2015
Last Updated:
11/25/2024

Operations

Publications

Comin M, Verzotto D. Beyond Fixed-Resolution Alignment-Free Measures for Mammalian Enhancers Sequence Comparison. IEEE/ACM Transactions on Computational Biology and Bioinformatics. 2014;11(4):628-637. doi:10.1109/tcbb.2014.2306830. PMID:26356333.

Comin M, Schimd M. Assembly-free genome comparison based on next-generation sequencing reads and variable length patterns. BMC Bioinformatics. 2014;15(S9). doi:10.1186/1471-2105-15-s9-s1. PMID:25252700. PMCID:PMC4168702.

Documentation

Links