seqbias
seqbias corrects protocol-specific sequence biases in high-throughput sequencing reads to improve accuracy of sequence abundance estimates for applications such as RNA-Seq de novo gene annotation and transcript quantification.
Key Features:
- Bias Quantification and Correction: Uses a graphical model to measure and correct sequence biases arising from protocol factors such as polymerase chain reaction amplification and primer affinities.
- Graphical Model Approach: Implements a Bayesian network trained on aligned reads and a reference genome and operates without requiring pre-existing gene annotations.
- Automatic Model Selection: Incorporates automatic model selection to determine model structure or parameters from the data.
- Minimal Risk of Spurious Adjustment: Designed to have negligible effects on unbiased data to avoid introducing false corrections.
- Empirical Validation: Demonstrated across multiple datasets to reduce bias and increase uniformity in sequence abundance measurements.
Scientific Applications:
- RNA-Seq transcript quantification: Improves accuracy of transcript abundance estimates from RNA-Seq data by correcting sequence-specific biases.
- De novo gene annotation: Supports more reliable inference of gene models from sequencing data by reducing protocol-induced bias in read distributions.
- General sequencing bias assessment: Enables evaluation and correction of protocol-specific biases in high-throughput sequencing experiments.
Methodology:
Implements a Bayesian network (graphical model) trained on aligned reads and a reference genome, with automatic model selection to learn and correct protocol-specific sequence biases without relying on gene annotations.
Topics
Collections
Details
- License:
- GPL-3.0
- Tool Type:
- command-line tool, library
- Operating Systems:
- Linux, Windows, Mac
- Programming Languages:
- R
- Added:
- 1/17/2017
- Last Updated:
- 12/30/2018
Operations
Data Inputs & Outputs
Sequence analysis
Publications
Jones DC, Ruzzo WL, Peng X, Katze MG. A new approach to bias correction in RNA-Seq. Bioinformatics. 2012;28(7):921-928. doi:10.1093/bioinformatics/bts055. PMID:22285831. PMCID:PMC3315719.