ECHO
ECHO corrects base-call errors in short-read sequencing data without a reference genome, improving read accuracy for downstream analyses such as de novo assembly and variant detection.
Key Features:
- Reference-free operation: Performs error correction without using a reference genome.
- Parameter optimization: Automatically sets parameters based on an assumed probabilistic model tailored to each sequencing run.
- Quality scoring: Assigns a quality score to each corrected base.
- Heterozygosity modeling: Explicitly models heterozygous sites in diploid genomes to identify bases originating from those sites.
- Coverage-aware performance: Improves error-correction accuracy across varying sequence coverage depths and particularly at the ends of reads.
- De novo assembly facilitation: Serves as a preprocessing step to improve input read quality for de novo genome assembly, especially under limited coverage.
Scientific Applications:
- Genome assembly: Provides cleaner short-read data to improve accuracy of de novo genome assembly.
- Variant calling: Enhances detection of variants by modeling heterozygous sites and correcting sequencing errors.
- Population genomics: Enables reference-free analyses across diverse populations or species and handles nonuniform coverage.
Methodology:
ECHO applies a reference-free error-correction algorithm that explicitly models heterozygous sites in diploid genomes, uses an assumed probabilistic model to automatically set parameters per sequencing run, and assigns quality scores to corrected bases.
Topics
Details
- License:
- BSD-3-Clause
- Tool Type:
- command-line tool
- Operating Systems:
- Linux, Windows, Mac
- Programming Languages:
- C++, Python
- Added:
- 1/13/2017
- Last Updated:
- 11/25/2024
Operations
Publications
Kao W, Chan AH, Song YS. ECHO: A reference-free short-read error correction algorithm. Genome Research. 2011;21(7):1181-1192. doi:10.1101/gr.111351.110. PMID:21482625. PMCID:PMC3129260.