TrieDedup

TrieDedup removes PCR duplicates from high-throughput sequencing reads using a trie-based algorithm that preserves accuracy when ambiguous bases ('N's) are present.


Key Features:

  • Efficient Deduplication: Achieves up to 270-fold faster deduplication of raw sequences compared to traditional pairwise comparison methods.
  • Handling Ambiguous Bases: Incorporates ambiguous bases ('N's) into the deduplication logic to accurately identify and remove PCR duplicates from low-quality reads.
  • Memory Optimization: Reduces memory usage by approximately 20% via a restrictedDict implementation in Python.
  • Versatile Applications: Supports PCR deduplication, barcode and Unique Molecular Identifier (UMI) assignment, and repertoire diversity analysis.

Scientific Applications:

  • PCR Duplicate Removal: Removes PCR duplicates from high-throughput sequencing datasets to improve data accuracy.
  • Barcode and UMI Assignment: Enables barcode and Unique Molecular Identifier (UMI) assignment workflows for large-scale sequencing analyses.
  • Repertoire Diversity Analysis: Facilitates repertoire diversity analysis for studies of genetic variability and function.

Methodology:

Sequences are stored and compared in a trie (prefix tree) that accounts for ambiguous bases ('N'); memory is optimized via a restrictedDict implemented in Python.

Topics

Details

License:
Apache-2.0
Cost:
Free of charge
Tool Type:
command-line tool
Programming Languages:
Python
Added:
6/18/2024
Last Updated:
11/24/2024

Operations

Publications

Hu J, Luo S, Tian M, Ye AY. TrieDedup: a fast trie-based deduplication algorithm to handle ambiguous bases in high-throughput sequencing. BMC Bioinformatics. 2024;25(1). doi:10.1186/s12859-024-05775-w. PMID:38637756. PMCID:PMC11025179.

PMID: 38637756
Funding: - National Institutes of Health: 5P01 AI138211-04, R01AI020047 - Bill & Melinda Gates Foundation Investment, United States: INV-021989