ErrorX
ErrorX corrects sequencing errors in next-generation sequencing (NGS) datasets to improve the accuracy of B- and T-cell receptor sequence analyses.
Key Features:
- Automated Error Correction: Employs deep learning to identify and correct bases with a high probability of being erroneous, including errors arising from PCR during library preparation and miscalled bases on sequencing instruments.
- Deep Learning Integration: Leverages machine learning models to distinguish true biological sequence variation from sequencing errors in B- and T-cell receptor data.
- Benchmark Performance: Demonstrated reductions in overall error rate in public datasets by up to 36% while maintaining a false positive rate of 0.05% or less.
- Universal Application: Applies directly to existing antibody and T-cell receptor sequencing datasets without requiring changes to library preparation protocols.
Scientific Applications:
- Immune repertoire profiling: Improves the reliability of deep profiling analyses of B-cell and T-cell receptor repertoires by reducing sequencing-induced artifacts.
- Immunology research and related studies: Enhances the accuracy of downstream biological conclusions drawn from NGS-based receptor sequencing datasets.
Methodology:
Trains deep learning models on known datasets to recognize patterns indicative of sequencing errors and to predict and correct erroneous bases.
Topics
Details
- Tool Type:
- command-line tool, desktop application
- Operating Systems:
- Mac, Linux, Windows
- Added:
- 1/18/2021
- Last Updated:
- 3/8/2021
Operations
Publications
Sevy AM. ErrorX: automated error correction for immune repertoire sequencing datasets. Unknown Journal. 2020. doi:10.1101/2020.02.17.952408.
Downloads
- Software packagehttps://github.com/EndeavorBio/ErrorX/releases
- Software packagehttps://github.com/EndeavorBio/ErrorX-Viewer/releases