MISHIMA

MISHIMA identifies shared rare oligonucleotide sequences to enable scalable multiple sequence alignment of large nucleotide datasets such as complete bacterial genomes for comparative genomics and evolutionary analysis.


Key Features:

  • Rare oligonucleotide identification: Identifies rare oligonucleotide sequences shared across all input sequences and uses them as anchors.
  • Avoidance of pairwise comparisons: Circumvents direct pairwise or progressive sequence comparisons to reduce computational complexity.
  • Divide and conquer segmentation: Segments sequences into smaller fragments using a divide and conquer strategy based on the identified markers.
  • External fragment alignment: Aligns resulting fragments independently using an external alignment program.
  • Assembly of partial alignments: Assembles partial fragment alignments into a comprehensive multiple sequence alignment.
  • Genome-scale handling: Designed to handle genome-scale nucleotide sequences, including complete bacterial genomes.
  • Input and output formats: Accepts FASTA input and produces MISHIMA or CLUSTALW formatted output.
  • Demonstrated performance: Demonstrated alignment of six complete Helicobacter pylori genomes (~1.7 Mb each) in approximately six hours on a single personal computer.

Scientific Applications:

  • Comparative genomics: Produces genome-scale multiple sequence alignments to support comparative analyses of bacterial genomes.
  • Evolutionary studies: Generates alignments suitable for genome-scale evolutionary analysis.

Methodology:

Identifies rare oligonucleotide sequences shared across all input sequences; uses these markers to segment sequences via a divide and conquer strategy; aligns fragments independently with an external alignment program; and assembles the partial alignments to construct the full multiple sequence alignment while avoiding direct pairwise or progressive alignments.

Topics

Details

Tool Type:
command-line tool
Operating Systems:
Windows
Programming Languages:
C
Added:
12/18/2017
Last Updated:
11/25/2024

Operations

Publications

Kryukov K, Saitou N. MISHIMA - a new method for high speed multiple alignment of nucleotide sequences of bacterial genome scale data. BMC Bioinformatics. 2010;11(1). doi:10.1186/1471-2105-11-142. PMID:20298584. PMCID:PMC2848238.

Documentation

Links