PBSIM2

PBSIM2 simulates long-read sequencing datasets from a reference genome, modeling quality scores and non-uniform error patterns characteristic of PacBio and Oxford Nanopore technologies.


Key Features:

  • Generative model for quality scores: Uses a hidden Markov Model (HMM) with factorized information criteria to capture and simulate non-uniform error characteristics of long-read sequencing.
  • Efficiency and scalability: Generates simulated datasets in minutes at high coverage (e.g., 100x) for large references (~10 million base pairs) with runtime scaling proportionally with coverage depth and reference length.
  • Memory requirements: Requires memory roughly equal to the length of the reference sequence plus a few megabytes.
  • Validation and consistency: Produces simulated reads that closely mimic real long-read sequencing data for reliable downstream analyses.

Scientific Applications:

  • Pipeline development and testing: Produces realistic datasets for developing and benchmarking bioinformatics pipelines for long-read data.
  • Sequence assembly: Enables evaluation and optimization of assembly algorithms under long-read error profiles.
  • Variant calling and other genomic analyses: Facilitates testing of variant calling and other analyses that require accurate modeling of read errors.

Methodology:

Simulates sequencing reads from a reference genome while modeling quality scores using a hidden Markov Model combined with factorized information criteria to represent observed non-uniform error patterns.

Topics

Details

License:
GPL-2.0
Programming Languages:
C++, Shell
Added:
1/18/2021
Last Updated:
11/24/2024

Operations

Publications

Ono Y, Asai K, Hamada M. PBSIM2: a simulator for long-read sequencers with a novel generative model of quality scores. Bioinformatics. 2020;37(5):589-595. doi:10.1093/bioinformatics/btaa835. PMID:32976553. PMCID:PMC8097687.

PMID: 32976553
PMCID: PMC8097687
Funding: - MEXT KAKENHI: JP16H05879, JP16H06279, JP20H00624, JP24680031, JP25240044