MISHIMA
MISHIMA identifies shared rare oligonucleotide sequences to enable scalable multiple sequence alignment of large nucleotide datasets such as complete bacterial genomes for comparative genomics and evolutionary analysis.
Key Features:
- Rare oligonucleotide identification: Identifies rare oligonucleotide sequences shared across all input sequences and uses them as anchors.
- Avoidance of pairwise comparisons: Circumvents direct pairwise or progressive sequence comparisons to reduce computational complexity.
- Divide and conquer segmentation: Segments sequences into smaller fragments using a divide and conquer strategy based on the identified markers.
- External fragment alignment: Aligns resulting fragments independently using an external alignment program.
- Assembly of partial alignments: Assembles partial fragment alignments into a comprehensive multiple sequence alignment.
- Genome-scale handling: Designed to handle genome-scale nucleotide sequences, including complete bacterial genomes.
- Input and output formats: Accepts FASTA input and produces MISHIMA or CLUSTALW formatted output.
- Demonstrated performance: Demonstrated alignment of six complete Helicobacter pylori genomes (~1.7 Mb each) in approximately six hours on a single personal computer.
Scientific Applications:
- Comparative genomics: Produces genome-scale multiple sequence alignments to support comparative analyses of bacterial genomes.
- Evolutionary studies: Generates alignments suitable for genome-scale evolutionary analysis.
Methodology:
Identifies rare oligonucleotide sequences shared across all input sequences; uses these markers to segment sequences via a divide and conquer strategy; aligns fragments independently with an external alignment program; and assembles the partial alignments to construct the full multiple sequence alignment while avoiding direct pairwise or progressive alignments.
Topics
Details
- Tool Type:
- command-line tool
- Operating Systems:
- Windows
- Programming Languages:
- C
- Added:
- 12/18/2017
- Last Updated:
- 11/25/2024
Operations
Publications
Kryukov K, Saitou N. MISHIMA - a new method for high speed multiple alignment of nucleotide sequences of bacterial genome scale data. BMC Bioinformatics. 2010;11(1). doi:10.1186/1471-2105-11-142. PMID:20298584. PMCID:PMC2848238.