Tapir

Tapir CBS performs reference-free identification of candidate reference DNAs by querying a server-hosted index with samples of sequencing reads to support organism identification and downstream analyses in contexts such as infectious disease, food safety, bioreactor management, and environmental monitoring.


Key Features:

  • Reference-Free Analysis: Identifies potential reference genomes for unidentified sequencing reads without requiring a predefined reference.
  • Distributed Architecture: Uses a client–server model where the client sends a sample of raw reads to a server that hosts an index of known reference DNAs and returns matching candidates.
  • Efficient Data Transfer: Transfers only a small subset of reads between client and server to minimize bandwidth usage.
  • Indexed Public Genomes: Server indexes tens of thousands of publicly available genomes and genomic regions from diverse organisms for candidate matching.
  • Client Implementations: Offers two client implementations: a browser-accessible client and a Python script for dataset processing.
  • Automated Processing and Quality Check: Supports automated sequencing data processing and provides instant quality checks for sequencing runs from desktop sequencers.

Scientific Applications:

  • Infectious Disease Surveillance: Tracking pathogens in clinical or environmental samples by identifying matching reference sequences.
  • Food Safety Assurance: Monitoring food products for contamination by pathogenic microorganisms through reference matching.
  • Bioreactor Optimization: Analyzing microbial communities within bioreactors by identifying constituent organisms from sequencing reads.
  • Environmental Studies: Identifying and cataloging biodiversity from complex environmental samples using reference candidates returned by the server.

Methodology:

The client sends a sample of unidentified sequencing reads to the server, the server returns a list of potential matching reference sequences, and those reference sequences can be downloaded for further computational analyses such as alignment.

Topics

Details

Maturity:
Emerging
Cost:
Free of charge (with restrictions)
Tool Type:
web application
Operating Systems:
Linux, Windows, Mac
Added:
1/21/2015
Last Updated:
11/25/2024

Operations

Publications

Gautier L, Lund O. Low-Bandwidth and Non-Compute Intensive Remote Identification of Microbes from Raw Sequencing Reads. PLoS ONE. 2013;8(12):e83784. doi:10.1371/journal.pone.0083784. PMID:24391826. PMCID:PMC3877093.

Documentation

Links