TagCleaner

TagCleaner removes WTA and MID tag sequences from primer-amplified metagenomic sequencing reads to improve the quality of data for downstream analyses such as taxonomic profiling and functional annotation.


Key Features:

  • Automatic tag detection and removal: Identifies known and unknown tag sequences (including WTA and MID tags) and removes them from reads while accommodating insertions and deletions.
  • Handling ambiguities: Accounts for ambiguous bases in primer sequences and for unavailable or incorrectly reported tag sequences during detection and trimming.
  • Data quality enhancement: Filters trimmed reads to remove duplicates, short reads, and reads with high rates of ambiguous bases, and splits artificial concatenated sequences arising from fragment-to-fragment concatenations.
  • Customizable parameters: Allows modification of filter parameters to adjust trimming and filtering thresholds.

Scientific Applications:

  • Taxonomic profiling: Produces cleaned reads suitable for accurate taxonomic assignment from metagenomic datasets.
  • Functional annotation: Improves sequence quality for reliable gene and function identification in metagenomes.
  • Comparative metagenomics: Enables reproducible comparisons across samples by removing amplification- and tag-derived artifacts.

Methodology:

Automatic identification of tag sequences, removal accommodating insertions and deletions, followed by filters that trim reads, remove duplicates and short/ambiguous reads, and split artificial concatenated sequences.

Topics

Details

License:
GPL-3.0
Maturity:
Mature
Tool Type:
web application
Operating Systems:
Linux, Windows, Mac
Programming Languages:
Perl
Added:
1/13/2017
Last Updated:
11/25/2024

Operations

Publications

Schmieder R, Lim YW, Rohwer F, Edwards R. TagCleaner: Identification and removal of tag sequences from genomic and metagenomic datasets. BMC Bioinformatics. 2010;11(1). doi:10.1186/1471-2105-11-341. PMID:20573248. PMCID:PMC2910026.

Documentation