VecScreen_plus_taxonomy

VecScreen_plus_taxonomy identifies and annotates vector contamination in nucleic acid sequences by using taxonomy-aware matching to support submission screening and curation in GenBank.


Key Features:

  • Taxonomy-Informed Analysis: Incorporates taxonomic information from the UniVec database to distinguish true-positive from false-positive VecScreen matches when sequences and matching vector subsequences derive from related organisms.
  • Retrospective Analysis Pipeline: A retrospective pipeline analyzed existing GenBank entries and resulted in identification and correction of over 8,000 contaminated sequences in the nonredundant nucleotide database.
  • Prospective Analysis Pipeline: A prospective pipeline, implemented since April 2017, evaluates new sequence submissions to GenBank for vector contamination.
  • Data Integration: Leverages source annotations from UniVec release 10.0 to assess the taxonomic origin of UniVec entries during contamination assessment.

Scientific Applications:

  • Database curation: Improves quality of public sequence repositories such as GenBank and the nonredundant nucleotide database by detecting and annotating vector contamination.
  • Submission quality control: Provides screening for new sequence submissions to identify vector-derived artifacts prior to accessioning.
  • Research data integrity: Reduces artifacts that could compromise experimental results or downstream analyses in genomics and bioinformatics studies.

Methodology:

VecScreen_plus_taxonomy runs VecScreen against the UniVec database (including UniVec release 10.0 source annotations) and classifies matches using taxonomy-aware rules within retrospective and prospective software pipelines.

Topics

Details

Tool Type:
command-line tool
Operating Systems:
Linux
Programming Languages:
Perl
Added:
6/21/2018
Last Updated:
11/25/2024

Operations

Publications

Schäffer AA, Nawrocki EP, Choi Y, Kitts PA, Karsch-Mizrachi I, McVeigh R. VecScreen_plus_taxonomy: imposing a tax(onomy) increase on vector contamination screening. Bioinformatics. 2017;34(5):755-759. doi:10.1093/bioinformatics/btx669. PMID:29069347. PMCID:PMC6030928.

Documentation