UnderIICRMS
UnderIICRMS performs alignment-free comparisons of cis-regulatory modules (CRMs) and unassembled genomes using variable-length patterns to detect regulatory sequence similarity and compare NGS read datasets.
Key Features:
- Parameter-Free Approach: Operates without predefined parameters, removing the need for parameter tuning across datasets and biological contexts.
- Variable-Length Patterns: Utilizes variable-length words/patterns instead of fixed-length k-mers to capture length variability in CRMs and genomic sequences.
- Enhancer Clustering and Functional Correlation: Identifies functionally similar enhancers and clusters within cis-regulatory modules by employing a similarity measure that detects related regulatory sequences.
- Discrimination of Functionally Related Enhancers: Distinguishes enhancers active in different tissues, as demonstrated on simulated data and ChIP-seq sequences including mouse enhancers.
- Genome Comparison Without Assembly: Enables direct comparison of unassembled genomes using NGS reads when reference genomes are unavailable.
- Similarity Measure Incorporating Reverses and Reverse-Complements: Defines a similarity measure based on variable-length patterns including reverses and reverse-complements and their statistical and syntactical properties.
- Performance in Genome Discrimination: Evaluated against other alignment-free statistics on simulated and real genomes and shown to outperform fixed-length pattern methods, particularly on real genome NGS read data.
Scientific Applications:
- Comparative Genomic Studies: Facilitates assembly-free comparison of genomes from NGS reads, applicable to non-model organisms and newly sequenced species.
- Regulatory Sequence Analysis: Supports analysis of enhancers and CRM architecture to infer functional correlations among regulatory elements.
- Evolutionary Biology: Enables study of evolutionary relationships between genomes using alignment-free comparisons of NGS data in the absence of a reference genome.
Methodology:
Performs alignment-free sequence comparisons using variable-length patterns/words in a parameter-free framework; computes a similarity measure that includes reverses and reverse-complements and assesses statistical and syntactical properties; validated on simulated data and ChIP-seq sequences and compared against other alignment-free statistics using unassembled NGS reads.
Topics
Details
- License:
- Other
- Tool Type:
- command-line tool
- Operating Systems:
- Linux, Windows, Mac
- Programming Languages:
- Java
- Added:
- 1/22/2015
- Last Updated:
- 11/25/2024
Operations
Publications
Comin M, Verzotto D. Beyond Fixed-Resolution Alignment-Free Measures for Mammalian Enhancers Sequence Comparison. IEEE/ACM Transactions on Computational Biology and Bioinformatics. 2014;11(4):628-637. doi:10.1109/tcbb.2014.2306830. PMID:26356333.
Comin M, Schimd M. Assembly-free genome comparison based on next-generation sequencing reads and variable length patterns. BMC Bioinformatics. 2014;15(S9). doi:10.1186/1471-2105-15-s9-s1. PMID:25252700. PMCID:PMC4168702.