TagDust

TagDust identifies and filters artifactual sequences generated during library preparation in high-throughput sequencing to improve the accuracy of downstream genomic analyses.


Key Features:

  • Artifact Identification: Compares reads to known sequences used in library preparation and detects artifacts explained by combinations or partial matches to those reference sequences.
  • User-Defined False Discovery Rate (FDR): Applies a user-specified FDR cutoff to control the stringency of artifact calling and balance sensitivity and specificity.
  • Application on High-Throughput Platforms: Demonstrated compatibility with data generated on Illumina's Genome Analyzer.

Scientific Applications:

  • Quality control: Identifies and removes artifactual reads to reduce background noise in sequencing libraries.
  • Gene expression profiling and variant calling: Improves the accuracy of downstream analyses such as gene expression quantification and variant detection by removing artifactual sequences.
  • Sequencing assay development: Aids assay validation by minimizing library-preparation–derived artifacts during development of novel sequencing assays.

Methodology:

Compares sequenced reads against a database of known library-preparation sequences, identifies reads explainable as combinations or partial matches to those references, flags them as potential artifacts, and applies a user-defined FDR cutoff to set calling stringency.

Topics

Collections

Details

License:
GPL-3.0
Maturity:
Mature
Tool Type:
command-line tool
Operating Systems:
Linux
Programming Languages:
C
Added:
1/13/2017
Last Updated:
11/25/2024

Operations

Publications

Lassmann T, Hayashizaki Y, Daub CO. TagDust—a program to eliminate artifacts from next generation sequencing data. Bioinformatics. 2009;25(21):2839-2840. doi:10.1093/bioinformatics/btp527. PMID:19737799. PMCID:PMC2781754.