thinspace
thinspace reduces large microbial genome collections by performing taxonomy-free dereplication using protein-content embeddings that represent bacterial and archaeal genomes as low-dimensional vectors for downstream metagenomic analyses.
Key Features:
- Protein-content embeddings: Represents genomes as low-dimensional numeric vectors derived from protein content using the nanotext concept.
- Taxonomy-free dereplication: Performs dereplication and grouping of genomes without relying on taxonomic labels or manual curation, using sequence data alone.
- Vector-based grouping: Groups and dereplicates similar genomes based on vector similarity of the embeddings.
- Frugal dereplication: Reduces large genome collections substantially (e.g., GTDB from ~150,000 to <22,000 genomes) while preserving representative diversity.
- Improved metagenomic classification: Increases the percentage of classified metagenomic reads (reported fivefold improvement versus NCBI RefSeq in the described comparison).
- Comparable to curated subsets: Achieves results comparable to larger and manually curated GTDB subsets while reducing dataset size.
- Resource efficiency: Enables dereplication on regular hardware without compromising downstream analytical outcomes.
Scientific Applications:
- Genome collection reduction: Creating compact, representative subsets of large microbial genome databases such as GTDB for downstream analyses.
- Metagenomic read classification: Improving the fraction of classified reads in metagenomic datasets by using reduced, representative reference sets.
- Taxonomy-free analysis: Enabling dereplication and representative selection in datasets where taxonomic labels are incomplete, biased, or absent.
- Reference database optimization: Producing smaller reference collections that maintain analytical utility for tasks like taxon assignment and comparative genomics.
Methodology:
Compute protein-content embeddings with the nanotext approach to represent genomes as low-dimensional vectors, then group and dereplicate similar genomes based on vector similarity using sequence-only data.
Topics
Details
- License:
- BSD-3-Clause
- Programming Languages:
- Python
- Added:
- 1/14/2020
- Last Updated:
- 1/16/2021
Operations
Publications
Viehweger A, Hoelzer M, Brandt C. Addressing dereplication crisis: Taxonomy-free reduction of massive genome collections using embeddings of protein content. Unknown Journal. 2019. doi:10.1101/855262.
DOI: 10.1101/855262