KIMI

KIMI identifies sequence motifs in metagenomic reads and controls the false discovery rate using model-X Knockoffs to select biologically relevant k-mers for distinguishing microbial groups such as bacteria and viruses.


Key Features:

  • False Discovery Rate Control: Uses model-X Knockoffs to control the false discovery rate (FDR) in motif selection, providing formal FDR guarantees.
  • High Power and Accuracy: Demonstrated in simulation studies to maintain higher power while controlling FDR compared with the Benjamini–Hochberg procedure and the q-value method.
  • Relevance-Based k-mer Selection: Selectively identifies k-mers relevant for distinguishing sequence groups instead of using all possible k-mers, concentrating on biologically informative subsequences of length k.
  • Application in Metagenomics: Targets motif discovery in metagenomic datasets to identify motifs specific to members of microbial communities.
  • Viral motif discovery example: Applied to datasets of viral and bacterial contigs where models trained on KIMI-selected k-mers increased accuracy of viral versus bacterial classification.

Scientific Applications:

  • Metagenomic motif discovery: Identification of sequence motifs in metagenomic reads to reveal taxa-specific patterns.
  • Microbial community classification: Feature selection for classifiers distinguishing bacteria, viruses, and other microbial groups.
  • Virology and microbiology research: Extraction of biologically relevant k-mers to support studies in microbiology, virology, and comparative genomics.

Methodology:

Uses model-X Knockoffs to perform motif discovery and control FDR, selects relevant k-mers from molecular sequences with an adjustable target FDR, and evaluates performance via simulation studies and comparison to the Benjamini–Hochberg and q-value methods.

Topics

Details

Tool Type:
library
Programming Languages:
C++, R
Added:
1/18/2021
Last Updated:
11/24/2024

Operations

Publications

Bai X, Ren J, Fan Y, Sun F. KIMI: Knockoff Inference for Motif Identification from molecular sequences with controlled false discovery rate. Bioinformatics. 2020;37(6):759-766. doi:10.1093/bioinformatics/btaa912. PMID:33119059. PMCID:PMC8599924.

PMID: 33119059
PMCID: PMC8599924
Funding: - US National Institutes of Health: 1R01GM131407, R01GM120624