SeqTrim

SeqTrim preprocesses sequence reads from Sanger sequencing and next-generation platforms such as pyrosequencing to identify and remove sequencing artefacts and improve sequence quality for downstream analyses.


Key Features:

  • Customizable Preprocessing Algorithms: Incorporates both previously published and newly designed algorithms tailored to identify and remove sequence inserts, low-quality reads, vector sequences, adaptor sequences, low-complexity regions, contaminant sequences, and chimeric reads.
  • Broad Compatibility with Input/Output Formats: Supports multiple input and output formats to facilitate incorporation into existing sequence processing pipelines.
  • Library and Read-Type Support: Efficiently handles sequences from EST libraries, SSH libraries, genomic DNA libraries, and pyrosequencing reads without leading to over-trimming.
  • Performance Excellence: Reveals more information per sequence compared to previously described preprocessors and improves discrimination between valid reads and artefacts.
  • Comprehensive Artefact Discardment: Implements a recommended pipeline that enables effective discarding of sequencing or experimental artefacts while preserving valuable sequence information.

Scientific Applications:

  • Genomics: Produces higher-quality genomic sequence data by removing vector, adaptor, contaminant, low-complexity, and chimeric sequences from genomic DNA libraries.
  • Transcriptomics: Preprocesses EST and SSH library reads to improve transcriptome assembly and downstream expression analyses.
  • Metagenomics: Reduces artefacts and contaminant sequences in high-throughput metagenomic sequencing datasets to improve community profiling.
  • Sequence Database Curation: Helps prevent database contamination by identifying and discarding artefactual sequences prior to deposition.

Methodology:

SeqTrim applies previously published and newly designed algorithms in a recommended, configurable preprocessing pipeline to detect and remove sequence inserts, low-quality regions, vector and adaptor sequences, low-complexity regions, contaminants, and chimeric reads.

Topics

Details

Tool Type:
web application
Operating Systems:
Linux, Windows, Mac
Programming Languages:
Perl
Added:
1/13/2017
Last Updated:
11/25/2024

Operations

Publications

Falgueras J, Lara AJ, Fernández-Pozo N, Cantón FR, Pérez-Trabado G, Claros MG. SeqTrim: a high-throughput pipeline for pre-processing any type of sequence read. BMC Bioinformatics. 2010;11(1). doi:10.1186/1471-2105-11-38. PMID:20089148. PMCID:PMC2832897.

Documentation