GenoTan
GenoTan identifies length variations in microsatellite alleles from short sequence reads to distinguish true allelic length variants from sequencing- and PCR-induced indel noise.
Key Features:
- Microsatellite genotyping: Accurately identifies length variations in microsatellite alleles using short sequence reads.
- Discretized Gaussian mixture model: Models allele length distributions with a discretized Gaussian mixture model to support precise genotype calls.
- Rules-based approach: Integrates a rules-based component with the mixture model to separate true length variants from noise.
- Bidirectional indel error modeling: Explicitly accounts for the bidirectional nature of insertion/deletion (indel) errors to improve genotyping accuracy.
- Homopolymer decomposition: Implements a homopolymer decomposition method that estimates bias toward insertion or deletion within homopolymer runs.
- Noise source modeling: Models common error sources including PCR amplification errors, individual cell mutations, and misalignment/mis-mapping.
- Performance: Validated on simulated data at 40× coverage with 94.9% genotyping accuracy and reported to be 5–30× faster than competing programs on the same dataset.
- Real-data validation: Evaluated on mixed sequence data from two inbred Drosophila lines and produced the highest true-positive rate among evaluated programs.
Scientific Applications:
- Microsatellite genotyping from short-read sequencing: Enables precise allele length calls for microsatellite loci using short sequencing reads.
- Genotyping in repetitive and homopolymer regions: Targets loci prone to homopolymer-associated indel errors and repetitive-sequence challenges.
- Discrimination of true variants from technical noise: Distinguishes genuine length polymorphisms from PCR- and sequencing-induced indels.
- Analysis of mixed or inbred samples: Applicable to mixed-sample sequencing and inbred-line comparisons, exemplified by Drosophila inbred-line data.
Methodology:
Integrates a discretized Gaussian mixture model with a rules-based approach, incorporates a homopolymer decomposition method to estimate insertion/deletion bias, and explicitly models bidirectional indel errors; validated using simulated 40× coverage datasets and mixed Drosophila inbred-line sequence data.
Topics
Details
- Tool Type:
- command-line tool
- Operating Systems:
- Linux
- Programming Languages:
- Perl
- Added:
- 8/3/2017
- Last Updated:
- 11/25/2024
Operations
Publications
Tae H, Kim D, McCormick J, Settlage RE, Garner HR. Discretized Gaussian mixture for genotyping of microsatellite loci containing homopolymer runs. Bioinformatics. 2013;30(5):652-659. doi:10.1093/bioinformatics/btt595. PMID:24135263. PMCID:PMC3933874.