VSEARCH

VSEARCH performs high-throughput nucleotide sequence processing and analysis for metagenomics, genomics, and population genomics, including searching, global alignment, clustering, dereplication, chimera detection, and FASTQ processing.


Key Features:

  • Multithreading and performance: 64-bit multithreaded application with vectorisation to leverage parallel processing for large datasets and alignment tasks.
  • Heuristic-based searching: Fast heuristic search based on shared words between query and target sequences.
  • Global sequence alignment: Full dynamic programming for optimal global sequence alignments.
  • Clustering and dereplication: Clustering by similarity with length pre-sorting, abundance pre-sorting, or user-defined order, and dereplication of full-length sequences or prefixes.
  • Chimera detection: Both reference-based and de novo chimera detection methods.
  • Pairwise alignment: Parallel computation of pairwise alignments.
  • FASTQ processing: FASTQ format detection, filtering, read quality statistics, merging of paired reads, shuffling, rereplication, and conversion between FASTQ formats.
  • Low-complexity masking: Masking of low-complexity sequences using the DUST algorithm.
  • Sequence operations: Reverse complementation, sorting, and subsampling of sequences.
  • Similarity definitions: Choice among various similarity definitions to tailor analyses.
  • Comparative performance: Demonstrates superior accuracy in searching, clustering, chimera detection, and subsampling compared to USEARCH, while being slower than USEARCH for clustering and chimera detection but faster for paired-end read merging and dereplication.

Scientific Applications:

  • Metagenomics: Accurate sequence processing, chimera detection, clustering, and paired-read merging for metagenomic studies.
  • Genomics: Nucleotide sequence alignment, clustering, and dereplication for genomics analyses.
  • Population genomics: Dereplication and subsampling to support population genomics workflows.
  • Paired-end read processing: Merging and downstream processing of paired reads, including quality filtering and format conversion.
  • Chimera detection and clustering: Reference-based and de novo chimera detection combined with similarity-based clustering for sequence-based analyses.

Methodology:

Methods explicitly include heuristic searching based on shared words, full dynamic programming for global alignment, clustering by similarity with length/abundance/user-defined sorting, full-length and prefix dereplication, reference-based and de novo chimera detection, parallel pairwise alignments, multithreading and vectorisation, FASTQ format detection and filtering, read quality statistics, paired-read merging, shuffling, rereplication, DUST low-complexity masking, reverse complementation, subsampling, and FASTQ format conversion.

Topics

Details

License:
GPL-3.0
Maturity:
Mature
Cost:
Free of charge
Tool Type:
command-line tool
Operating Systems:
Linux, Mac
Programming Languages:
C++
Added:
3/7/2016
Last Updated:
11/24/2024

Operations

Data Inputs & Outputs

Publications

Rognes T, Flouri T, Nichols B, Quince C, Mahé F. VSEARCH: a versatile open source tool for metagenomics. PeerJ. 2016;4:e2584. doi:10.7717/peerj.2584. PMID:27781170. PMCID:PMC5075697.

PMID: 27781170
PMCID: PMC5075697
Funding: - UNINETT Sigma2: NN9383K - MRC Cloud Infrastructure for Microbial Bioinformatics (CLIMB): MR/L015080/1, MR/M50161X/1 - Deutsche Forschungsgemeinschaft: #DU1319/1-1

Documentation

Citation instructions', 'General
https://github.com/torognes/vsearch

Links