GenoTan

GenoTan identifies length variations in microsatellite alleles from short sequence reads to distinguish true allelic length variants from sequencing- and PCR-induced indel noise.


Key Features:

  • Microsatellite genotyping: Accurately identifies length variations in microsatellite alleles using short sequence reads.
  • Discretized Gaussian mixture model: Models allele length distributions with a discretized Gaussian mixture model to support precise genotype calls.
  • Rules-based approach: Integrates a rules-based component with the mixture model to separate true length variants from noise.
  • Bidirectional indel error modeling: Explicitly accounts for the bidirectional nature of insertion/deletion (indel) errors to improve genotyping accuracy.
  • Homopolymer decomposition: Implements a homopolymer decomposition method that estimates bias toward insertion or deletion within homopolymer runs.
  • Noise source modeling: Models common error sources including PCR amplification errors, individual cell mutations, and misalignment/mis-mapping.
  • Performance: Validated on simulated data at 40× coverage with 94.9% genotyping accuracy and reported to be 5–30× faster than competing programs on the same dataset.
  • Real-data validation: Evaluated on mixed sequence data from two inbred Drosophila lines and produced the highest true-positive rate among evaluated programs.

Scientific Applications:

  • Microsatellite genotyping from short-read sequencing: Enables precise allele length calls for microsatellite loci using short sequencing reads.
  • Genotyping in repetitive and homopolymer regions: Targets loci prone to homopolymer-associated indel errors and repetitive-sequence challenges.
  • Discrimination of true variants from technical noise: Distinguishes genuine length polymorphisms from PCR- and sequencing-induced indels.
  • Analysis of mixed or inbred samples: Applicable to mixed-sample sequencing and inbred-line comparisons, exemplified by Drosophila inbred-line data.

Methodology:

Integrates a discretized Gaussian mixture model with a rules-based approach, incorporates a homopolymer decomposition method to estimate insertion/deletion bias, and explicitly models bidirectional indel errors; validated using simulated 40× coverage datasets and mixed Drosophila inbred-line sequence data.

Topics

Details

Tool Type:
command-line tool
Operating Systems:
Linux
Programming Languages:
Perl
Added:
8/3/2017
Last Updated:
11/25/2024

Operations

Publications

Tae H, Kim D, McCormick J, Settlage RE, Garner HR. Discretized Gaussian mixture for genotyping of microsatellite loci containing homopolymer runs. Bioinformatics. 2013;30(5):652-659. doi:10.1093/bioinformatics/btt595. PMID:24135263. PMCID:PMC3933874.

Documentation

Links