synggen

synggen generates synthetic cancer sequencing datasets for whole-exome and targeted sequencing to support benchmarking of computational methods in translational cancer genomics and precision medicine.


Key Features:

  • Fast and Scalable Data Generation: Produces large-scale synthetic datasets for whole-exome and targeted sequencing to meet high-throughput benchmarking needs.
  • Realistic and Heterogeneous Datasets: Simulates phased germline single nucleotide polymorphisms (SNPs) and complex allele-specific somatic events to mimic tumor heterogeneity.
  • Versatility Across Platforms: Generates datasets representative of distinct sequencing platforms to enable cross-platform evaluation of computational methods.

Scientific Applications:

  • Benchmarking Computational Tools: Provides realistic synthetic datasets for rigorous testing and comparison of somatic genomic profiling and other computational methods.
  • Precision Medicine Research: Supports translational cancer genomics and precision medicine studies that require controlled, complex datasets for method validation.

Methodology:

Employs algorithms to simulate phased germline SNPs and allele-specific somatic variations to generate large-scale synthetic whole-exome and targeted sequencing datasets.

Topics

Details

License:
MIT
Cost:
Free of charge
Tool Type:
command-line tool
Operating Systems:
Linux
Programming Languages:
C
Added:
2/10/2023
Last Updated:
11/24/2024

Operations

Data Inputs & Outputs

Publications

Scandino R, Calabrese F, Romanel A. Synggen: fast and data-driven generation of synthetic heterogeneous NGS cancer data. Bioinformatics. 2022;39(1). doi:10.1093/bioinformatics/btac792. PMID:36484701. PMCID:PMC9825741.