VSEARCH
VSEARCH performs high-throughput nucleotide sequence processing and analysis for metagenomics, genomics, and population genomics, including searching, global alignment, clustering, dereplication, chimera detection, and FASTQ processing.
Key Features:
- Multithreading and performance: 64-bit multithreaded application with vectorisation to leverage parallel processing for large datasets and alignment tasks.
- Heuristic-based searching: Fast heuristic search based on shared words between query and target sequences.
- Global sequence alignment: Full dynamic programming for optimal global sequence alignments.
- Clustering and dereplication: Clustering by similarity with length pre-sorting, abundance pre-sorting, or user-defined order, and dereplication of full-length sequences or prefixes.
- Chimera detection: Both reference-based and de novo chimera detection methods.
- Pairwise alignment: Parallel computation of pairwise alignments.
- FASTQ processing: FASTQ format detection, filtering, read quality statistics, merging of paired reads, shuffling, rereplication, and conversion between FASTQ formats.
- Low-complexity masking: Masking of low-complexity sequences using the DUST algorithm.
- Sequence operations: Reverse complementation, sorting, and subsampling of sequences.
- Similarity definitions: Choice among various similarity definitions to tailor analyses.
- Comparative performance: Demonstrates superior accuracy in searching, clustering, chimera detection, and subsampling compared to USEARCH, while being slower than USEARCH for clustering and chimera detection but faster for paired-end read merging and dereplication.
Scientific Applications:
- Metagenomics: Accurate sequence processing, chimera detection, clustering, and paired-read merging for metagenomic studies.
- Genomics: Nucleotide sequence alignment, clustering, and dereplication for genomics analyses.
- Population genomics: Dereplication and subsampling to support population genomics workflows.
- Paired-end read processing: Merging and downstream processing of paired reads, including quality filtering and format conversion.
- Chimera detection and clustering: Reference-based and de novo chimera detection combined with similarity-based clustering for sequence-based analyses.
Methodology:
Methods explicitly include heuristic searching based on shared words, full dynamic programming for global alignment, clustering by similarity with length/abundance/user-defined sorting, full-length and prefix dereplication, reference-based and de novo chimera detection, parallel pairwise alignments, multithreading and vectorisation, FASTQ format detection and filtering, read quality statistics, paired-read merging, shuffling, rereplication, DUST low-complexity masking, reverse complementation, subsampling, and FASTQ format conversion.
Topics
Details
- License:
- GPL-3.0
- Maturity:
- Mature
- Cost:
- Free of charge
- Tool Type:
- command-line tool
- Operating Systems:
- Linux, Mac
- Programming Languages:
- C++
- Added:
- 3/7/2016
- Last Updated:
- 11/24/2024
Operations
Data Inputs & Outputs
Chimera detection
Publications
Rognes T, Flouri T, Nichols B, Quince C, Mahé F. VSEARCH: a versatile open source tool for metagenomics. PeerJ. 2016;4:e2584. doi:10.7717/peerj.2584. PMID:27781170. PMCID:PMC5075697.