EnsembleClust
EnsembleClust performs hierarchical clustering of noncoding RNAs to group unannotated transcripts into putative families based on sequence and secondary-structure similarity.
Key Features:
- Innovative Similarity Measure: Leverages suboptimal solutions within dynamic programming frameworks to define a similarity measure that improves accuracy of approximate ncRNA clustering algorithms.
- Comprehensive Alignment Strategy: Considers all possible sequence alignments and secondary structures rather than a single optimal alignment to assess ncRNA similarity.
- Balanced Performance: Maintains clustering quality while controlling computational cost, enabling performance when sequence identity among family members falls below 60%.
- Fast and Accurate Clustering: Simplifies approximation of structural alignment to achieve faster clustering suitable for large-scale genomic datasets.
Scientific Applications:
- Discovery of novel ncRNA families: Facilitates identification of new noncoding RNA families by grouping unannotated transcripts based on sequence and structure.
- Evolutionary and functional inference: Supports inference of evolutionary relationships and potential functional similarity among ncRNAs with low primary sequence identity.
- Annotation of unannotated transcripts: Assists annotation efforts by clustering unannotated transcripts into candidate families for further analysis.
Methodology:
Uses approximate structural alignment by incorporating suboptimal dynamic-programming solutions and considering all possible sequence alignments and secondary structures, followed by hierarchical clustering.
Topics
Details
- Tool Type:
- command-line tool
- Operating Systems:
- Linux
- Programming Languages:
- Ruby, Perl, C
- Added:
- 12/18/2017
- Last Updated:
- 11/25/2024
Operations
Data Inputs & Outputs
Clustering
Outputs
Publications
Saito Y, Sato K, Sakakibara Y. Fast and accurate clustering of noncoding RNAs using ensembles of sequence alignments and secondary structures. BMC Bioinformatics. 2011;12(S1). doi:10.1186/1471-2105-12-s1-s48. PMID:21342580. PMCID:PMC3044305.