thinspace

thinspace reduces large microbial genome collections by performing taxonomy-free dereplication using protein-content embeddings that represent bacterial and archaeal genomes as low-dimensional vectors for downstream metagenomic analyses.


Key Features:

  • Protein-content embeddings: Represents genomes as low-dimensional numeric vectors derived from protein content using the nanotext concept.
  • Taxonomy-free dereplication: Performs dereplication and grouping of genomes without relying on taxonomic labels or manual curation, using sequence data alone.
  • Vector-based grouping: Groups and dereplicates similar genomes based on vector similarity of the embeddings.
  • Frugal dereplication: Reduces large genome collections substantially (e.g., GTDB from ~150,000 to <22,000 genomes) while preserving representative diversity.
  • Improved metagenomic classification: Increases the percentage of classified metagenomic reads (reported fivefold improvement versus NCBI RefSeq in the described comparison).
  • Comparable to curated subsets: Achieves results comparable to larger and manually curated GTDB subsets while reducing dataset size.
  • Resource efficiency: Enables dereplication on regular hardware without compromising downstream analytical outcomes.

Scientific Applications:

  • Genome collection reduction: Creating compact, representative subsets of large microbial genome databases such as GTDB for downstream analyses.
  • Metagenomic read classification: Improving the fraction of classified reads in metagenomic datasets by using reduced, representative reference sets.
  • Taxonomy-free analysis: Enabling dereplication and representative selection in datasets where taxonomic labels are incomplete, biased, or absent.
  • Reference database optimization: Producing smaller reference collections that maintain analytical utility for tasks like taxon assignment and comparative genomics.

Methodology:

Compute protein-content embeddings with the nanotext approach to represent genomes as low-dimensional vectors, then group and dereplicate similar genomes based on vector similarity using sequence-only data.

Topics

Details

License:
BSD-3-Clause
Programming Languages:
Python
Added:
1/14/2020
Last Updated:
1/16/2021

Operations

Publications

Viehweger A, Hoelzer M, Brandt C. Addressing dereplication crisis: Taxonomy-free reduction of massive genome collections using embeddings of protein content. Unknown Journal. 2019. doi:10.1101/855262.