GENome EXogenous (GENEX) sequence detection workflow
GENome EXogenous (GENEX) sequence detection workflow detects and maps microbial-like and human-like sequences within eukaryotic and prokaryotic reference genomes to identify exogenous sequence contamination that can bias taxonomic profiling of environmental and ancient DNA.
Key Features:
- Input and Output Formats: Accepts reference sequences in FASTA format and outputs coordinates of identified regions in BED format.
- Indexing and Alignment: Constructs a Bowtie2 index from the provided reference genome and aligns pre-computed pseudo-reads to the index.
- Reference Databases: Uses pseudo-reads derived from GTDB v.214 and NCBI RefSeq release 213 for microbial sequences and from hg38 for human sequences.
- Custom Detection Scripts: Employs custom scripts to detect positions of covered regions within references and to quantify the most abundant microbial species when scanning eukaryotic genomes for microbial-like sequences.
- Genome-scale Analysis: Applies the workflow to nearly 3,000 reference genomes from NCBI RefSeq and GenBank (vertebrates, invertebrates, and plants) and to 1,323 PhyloNorway plant genome assemblies.
- Annotations and Masking Support: Produces genomic coordinates and taxonomic annotations of microbial-like or human-like regions to enable downstream masking during profiling analyses.
Scientific Applications:
- Contamination detection in reference genomes: Identifies microbial-like and human-like sequence contamination within eukaryotic and prokaryotic reference genomes to reduce biases in downstream analyses.
- Metagenomic taxonomic profiling: Improves reliability of taxonomic assignments in metagenomic datasets by providing regions to mask that would otherwise produce spurious mappings.
- Ancient and environmental DNA studies: Supports profiling and reconstruction of past ecosystems from low-abundance plant and animal DNA in palaeontological, archaeological, and high-latitude/herbarium-derived samples.
Methodology:
Constructs a Bowtie2 index from each reference, aligns pre-computed pseudo-reads derived from GTDB v.214, NCBI RefSeq release 213, or hg38, uses custom scripts to detect covered region positions and quantify abundant microbial species, and outputs BED coordinates with taxonomic annotations after analysis of ~3,000 RefSeq/GenBank genomes and 1,323 PhyloNorway plant assemblies.
Details
- Added:
- 8/10/2025
- Last Updated:
- 8/10/2025
Operations
Publications
Oskolkov N, Jin C, Clinton SL, Guinet B, Wijnands F, Johnson E, Kutschera VE, Kinsella CM, Heintzman PD, van der Valk T. Disinfecting eukaryotic reference genomes to improve taxonomic inference from ancient environmental metagenomic data. Unknown Journal. 2025. doi:10.1101/2025.03.19.644176.