PBSIM2
PBSIM2 simulates long-read sequencing datasets from a reference genome, modeling quality scores and non-uniform error patterns characteristic of PacBio and Oxford Nanopore technologies.
Key Features:
- Generative model for quality scores: Uses a hidden Markov Model (HMM) with factorized information criteria to capture and simulate non-uniform error characteristics of long-read sequencing.
- Efficiency and scalability: Generates simulated datasets in minutes at high coverage (e.g., 100x) for large references (~10 million base pairs) with runtime scaling proportionally with coverage depth and reference length.
- Memory requirements: Requires memory roughly equal to the length of the reference sequence plus a few megabytes.
- Validation and consistency: Produces simulated reads that closely mimic real long-read sequencing data for reliable downstream analyses.
Scientific Applications:
- Pipeline development and testing: Produces realistic datasets for developing and benchmarking bioinformatics pipelines for long-read data.
- Sequence assembly: Enables evaluation and optimization of assembly algorithms under long-read error profiles.
- Variant calling and other genomic analyses: Facilitates testing of variant calling and other analyses that require accurate modeling of read errors.
Methodology:
Simulates sequencing reads from a reference genome while modeling quality scores using a hidden Markov Model combined with factorized information criteria to represent observed non-uniform error patterns.
Topics
Details
- License:
- GPL-2.0
- Programming Languages:
- C++, Shell
- Added:
- 1/18/2021
- Last Updated:
- 11/24/2024
Operations
Publications
Ono Y, Asai K, Hamada M. PBSIM2: a simulator for long-read sequencers with a novel generative model of quality scores. Bioinformatics. 2020;37(5):589-595. doi:10.1093/bioinformatics/btaa835. PMID:32976553. PMCID:PMC8097687.
PMID: 32976553
PMCID: PMC8097687
Funding: - MEXT KAKENHI: JP16H05879, JP16H06279, JP20H00624, JP24680031, JP25240044