BFC
BFC corrects sequencing errors in Illumina short-read datasets to improve base-level accuracy for high-coverage human whole-genome and smaller-genome analyses.
Key Features:
- Error Correction Efficiency: Corrects a large number of sequencing errors while reducing overcorrection and false positives.
- Systematic Error Suppression: Suppresses systematic sequencing errors common in Illumina datasets, improving base accuracy for downstream analyses including de novo genome assemblies.
- Algorithmic Approach: Implements a non-greedy algorithmic approach to error correction, offering thorough correction with computational efficiency and speed comparable to greedy methods.
- Performance on Real Data: Demonstrated superior error-correction performance on real Illumina datasets relative to existing tools.
Scientific Applications:
- Illumina data quality improvement: Enhances base-level quality of Illumina short-read sequencing datasets prior to downstream analyses.
- High-coverage human whole-genome sequencing: Serves as a preprocessing step for high-coverage human WGS projects to reduce sequencing errors.
- De novo assembly and smaller genomes: Improves input read accuracy for de novo genome assembly and analyses of smaller genomes by suppressing systematic errors.
Methodology:
Applies Bayesian filtering using a non-greedy algorithm that targets systematic Illumina error patterns to identify and correct errors in short reads.
Topics
Details
- Tool Type:
- command-line tool
- Operating Systems:
- Linux
- Programming Languages:
- JavaScript, C++, C
- Added:
- 8/3/2017
- Last Updated:
- 11/25/2024
Operations
Publications
Li H. BFC: correcting Illumina sequencing errors. Bioinformatics. 2015;31(17):2885-2887. doi:10.1093/bioinformatics/btv290. PMID:25953801. PMCID:PMC4635656.
Documentation
General
https://github.com/lh3/bfc