fqduplicate
fqduplicate identifies and filters duplicate reads in sequencing datasets by comparing read names and quality sum scores to retain the highest-quality instance of identical sequences.
Key Features:
- Duplicate Detection: Detects duplicate reads by identifying previously encountered read names corresponding to identical sequences.
- Quality Score Evaluation: Compares quality sum scores among duplicate reads to determine and retain the highest-quality instance.
- Duplicate Reporting: Produces a list of read names corresponding to reads classified as duplicates with inferior quality scores.
Scientific Applications:
- Data Quality Assurance: Removes lower-quality duplicate reads to reduce bias and improve the integrity of sequencing datasets.
- Efficient Data Processing: Reduces redundant low-quality reads to streamline downstream analyses such as variant calling and metagenomic profiling.
Methodology:
Scan sequencing data to identify previously encountered read names, compare their quality sum scores, retain the instance with the superior score, and output read names corresponding to duplicates with inferior scores.
Topics
Collections
Details
- Maturity:
- Mature
- Tool Type:
- web application
- Added:
- 12/19/2016
- Last Updated:
- 8/27/2025
Operations
Publications
Afgan E, Baker D, van den Beek M, Blankenberg D, Bouvier D, Čech M, Chilton J, Clements D, Coraor N, Eberhard C, Grüning B, Guerler A, Hillman-Jackson J, Von Kuster G, Rasche E, Soranzo N, Turaga N, Taylor J, Nekrutenko A, Goecks J. The Galaxy platform for accessible, reproducible and collaborative biomedical analyses: 2016 update. Nucleic Acids Research. 2016;44(W1):W3-W10. doi:10.1093/nar/gkw343. PMID:27137889. PMCID:PMC4987906.
Mareuil F, Doppelt-Azeroual O, Ménager H. A public Galaxy platform at Pasteur used as an execution engine for web services. Unknown Journal. 2017. doi:10.7490/f1000research.1114334.1.