TagDust
TagDust identifies and filters artifactual sequences generated during library preparation in high-throughput sequencing to improve the accuracy of downstream genomic analyses.
Key Features:
- Artifact Identification: Compares reads to known sequences used in library preparation and detects artifacts explained by combinations or partial matches to those reference sequences.
- User-Defined False Discovery Rate (FDR): Applies a user-specified FDR cutoff to control the stringency of artifact calling and balance sensitivity and specificity.
- Application on High-Throughput Platforms: Demonstrated compatibility with data generated on Illumina's Genome Analyzer.
Scientific Applications:
- Quality control: Identifies and removes artifactual reads to reduce background noise in sequencing libraries.
- Gene expression profiling and variant calling: Improves the accuracy of downstream analyses such as gene expression quantification and variant detection by removing artifactual sequences.
- Sequencing assay development: Aids assay validation by minimizing library-preparation–derived artifacts during development of novel sequencing assays.
Methodology:
Compares sequenced reads against a database of known library-preparation sequences, identifies reads explainable as combinations or partial matches to those references, flags them as potential artifacts, and applies a user-defined FDR cutoff to set calling stringency.
Topics
Collections
Details
- License:
- GPL-3.0
- Maturity:
- Mature
- Tool Type:
- command-line tool
- Operating Systems:
- Linux
- Programming Languages:
- C
- Added:
- 1/13/2017
- Last Updated:
- 11/25/2024
Operations
Publications
Lassmann T, Hayashizaki Y, Daub CO. TagDust—a program to eliminate artifacts from next generation sequencing data. Bioinformatics. 2009;25(21):2839-2840. doi:10.1093/bioinformatics/btp527. PMID:19737799. PMCID:PMC2781754.