gmap_iit_store

gmap_iit_store constructs a map store for known genes and single nucleotide polymorphisms (SNPs) and performs cDNA-to-genome mapping and alignment leveraging GMAP algorithms for genomic mapping and alignment tasks.


Key Features:

  • Map store creation: Constructs a map store for known genes and single nucleotide polymorphisms (SNPs).
  • cDNA-to-genome mapping and alignment: Maps and aligns cDNA sequences to genomes using GMAP algorithms.
  • Low startup time and memory requirements: Operates with minimal startup time and reduced memory usage for individual and batch processing.
  • Accurate gene structure identification: Generates precise gene structures in the presence of substantial polymorphisms and sequence errors without relying on probabilistic splice site models.
  • Minimal sampling strategy: Uses a minimal sampling strategy for efficient genomic mapping.
  • Oligomer chaining: Employs oligomer chaining to enable approximate alignment of sequences.
  • Sandwich Dynamic Programming: Applies sandwich dynamic programming for robust splice site detection.
  • Microexon identification with statistical significance testing: Identifies microexons and assesses their statistical significance.
  • High splice-site recovery on mutated mRNAs: Identified all splice sites in over 99.3% of human mRNA sequences with 1% and 3% random mutations.
  • Comparative alignment quality: Produced higher-quality alignments more frequently than BLAT on human expressed sequence tags and than GeneSeqer on Arabidopsis cDNAs.
  • Increased speed: Achieves a several-fold increase in speed compared with existing genomic mapping programs.

Scientific Applications:

  • Gene Expression Analysis: Facilitates study of gene expression by mapping cDNA sequences to genomes.
  • Polymorphism Studies: Enables examination of genetic variation such as single nucleotide polymorphisms (SNPs).
  • Transcriptome Research: Supports transcriptome analyses by aligning expressed sequence tags and cDNAs for transcript reconstruction.
  • Large-scale sequencing data processing: Applies to next-generation DNA sequencing datasets requiring accurate and rapid alignment at scale.

Methodology:

Leverages GMAP algorithms to map and align cDNA to genomes and construct a map store for genes and SNPs using a minimal sampling strategy, oligomer chaining, sandwich dynamic programming for splice-site detection, and statistical significance testing for microexon identification.

Topics

Collections

Details

Maturity:
Mature
Tool Type:
web application
Operating Systems:
Linux, Windows, Mac
Programming Languages:
C
Added:
12/19/2016
Last Updated:
11/25/2024

Operations

Data Inputs & Outputs

Publications

Wu TD, Watanabe CK. GMAP: a genomic mapping and alignment program for mRNA and EST sequences. Bioinformatics. 2005;21(9):1859-1875. doi:10.1093/bioinformatics/bti310. PMID:15728110.

Afgan E, Baker D, van den Beek M, Blankenberg D, Bouvier D, Čech M, Chilton J, Clements D, Coraor N, Eberhard C, Grüning B, Guerler A, Hillman-Jackson J, Von Kuster G, Rasche E, Soranzo N, Turaga N, Taylor J, Nekrutenko A, Goecks J. The Galaxy platform for accessible, reproducible and collaborative biomedical analyses: 2016 update. Nucleic Acids Research. 2016;44(W1):W3-W10. doi:10.1093/nar/gkw343. PMID:27137889. PMCID:PMC4987906.

Mareuil F, Doppelt-Azeroual O, Ménager H. A public Galaxy platform at Pasteur used as an execution engine for web services. Unknown Journal. 2017. doi:10.7490/f1000research.1114334.1.

Documentation

Links