gemerr

gemerr generates simulated next-generation sequencing reads using empirically derived, sequence-context-based error models to reproduce run- and technology-specific error profiles for analyses such as SNP calling.


Key Features:

  • Empirical Error Modeling: Derives sequence-context-based error profiles from real sequencing runs, including modeling of indels and quality score variations.
  • Format and Platform Support: Produces single- and paired-end reads compatible with SAM and FASTQ formats and reflects platforms such as Illumina and Roche/454.
  • Versatile Simulation Capabilities: Simulates reads from one or more genomes or haplotype sets to support diverse experimental scenarios.
  • Error Profile Comparison: Derives and compares error models from multiple sequencing runs to characterize run-to-run and technology-specific differences.
  • Impact on Variant Calling: Demonstrates effects of error profiles on SNP-calling accuracy, including findings that VarScan retains accuracy primarily for SNPs with frequency greater than 3% and that optimal VarScan 'minimum average quality' settings vary with error model.
  • Empirical Fragment and Quality Distributions: Generates reads reflecting observed fragment length and quality score distributions from actual sequencing experiments.

Scientific Applications:

  • Deep sequencing evaluation: Assess how sequencing error profiles affect detection of low-frequency variants in deep sequencing experiments.
  • Metagenomic studies: Simulate complex communities to evaluate how run-specific errors influence taxonomic and diversity analyses.
  • Resequencing projects: Benchmark variant-calling pipelines and parameter settings under realistic error conditions.
  • Variant-calling validation: Quantify the impact of different error models on SNP-calling accuracy and on parameters such as VarScan's 'minimum average quality'.
  • Genetic diversity assessment: Explore how sequencing-run variability in errors alters estimates of genetic diversity and low-frequency SNP detection.

Methodology:

Uses empirical sequencing run data to derive context-specific error models; generates reads that reflect empirical fragment length and quality score distributions; analyzes how differing error profiles affect SNP-calling accuracy and parameter sensitivity (e.g., VarScan 'minimum average quality').

Topics

Collections

Details

Maturity:
Mature
Tool Type:
web application
Operating Systems:
Linux, Windows, Mac
Programming Languages:
Python
Added:
12/19/2016
Last Updated:
11/25/2024

Operations

Data Inputs & Outputs

Modelling and simulation

Publications

McElroy KE, Luciani F, Thomas T. GemSIM: general, error-model based simulator of next-generation sequencing data. BMC Genomics. 2012;13(1):74. doi:10.1186/1471-2164-13-74. PMID:22336055. PMCID:PMC3305602.

Afgan E, Baker D, van den Beek M, Blankenberg D, Bouvier D, Čech M, Chilton J, Clements D, Coraor N, Eberhard C, Grüning B, Guerler A, Hillman-Jackson J, Von Kuster G, Rasche E, Soranzo N, Turaga N, Taylor J, Nekrutenko A, Goecks J. The Galaxy platform for accessible, reproducible and collaborative biomedical analyses: 2016 update. Nucleic Acids Research. 2016;44(W1):W3-W10. doi:10.1093/nar/gkw343. PMID:27137889. PMCID:PMC4987906.

Mareuil F, Doppelt-Azeroual O, Ménager H. A public Galaxy platform at Pasteur used as an execution engine for web services. Unknown Journal. 2017. doi:10.7490/f1000research.1114334.1.

Documentation

Links