TrieDedup
TrieDedup removes PCR duplicates from high-throughput sequencing reads using a trie-based algorithm that preserves accuracy when ambiguous bases ('N's) are present.
Key Features:
- Efficient Deduplication: Achieves up to 270-fold faster deduplication of raw sequences compared to traditional pairwise comparison methods.
- Handling Ambiguous Bases: Incorporates ambiguous bases ('N's) into the deduplication logic to accurately identify and remove PCR duplicates from low-quality reads.
- Memory Optimization: Reduces memory usage by approximately 20% via a restrictedDict implementation in Python.
- Versatile Applications: Supports PCR deduplication, barcode and Unique Molecular Identifier (UMI) assignment, and repertoire diversity analysis.
Scientific Applications:
- PCR Duplicate Removal: Removes PCR duplicates from high-throughput sequencing datasets to improve data accuracy.
- Barcode and UMI Assignment: Enables barcode and Unique Molecular Identifier (UMI) assignment workflows for large-scale sequencing analyses.
- Repertoire Diversity Analysis: Facilitates repertoire diversity analysis for studies of genetic variability and function.
Methodology:
Sequences are stored and compared in a trie (prefix tree) that accounts for ambiguous bases ('N'); memory is optimized via a restrictedDict implemented in Python.
Topics
Details
- License:
- Apache-2.0
- Cost:
- Free of charge
- Tool Type:
- command-line tool
- Programming Languages:
- Python
- Added:
- 6/18/2024
- Last Updated:
- 11/24/2024
Operations
Publications
Hu J, Luo S, Tian M, Ye AY. TrieDedup: a fast trie-based deduplication algorithm to handle ambiguous bases in high-throughput sequencing. BMC Bioinformatics. 2024;25(1). doi:10.1186/s12859-024-05775-w. PMID:38637756. PMCID:PMC11025179.
PMID: 38637756
PMCID: PMC11025179
Funding: - National Institutes of Health: 5P01 AI138211-04, R01AI020047
- Bill & Melinda Gates Foundation Investment, United States: INV-021989