KIMI
KIMI identifies sequence motifs in metagenomic reads and controls the false discovery rate using model-X Knockoffs to select biologically relevant k-mers for distinguishing microbial groups such as bacteria and viruses.
Key Features:
- False Discovery Rate Control: Uses model-X Knockoffs to control the false discovery rate (FDR) in motif selection, providing formal FDR guarantees.
- High Power and Accuracy: Demonstrated in simulation studies to maintain higher power while controlling FDR compared with the Benjamini–Hochberg procedure and the q-value method.
- Relevance-Based k-mer Selection: Selectively identifies k-mers relevant for distinguishing sequence groups instead of using all possible k-mers, concentrating on biologically informative subsequences of length k.
- Application in Metagenomics: Targets motif discovery in metagenomic datasets to identify motifs specific to members of microbial communities.
- Viral motif discovery example: Applied to datasets of viral and bacterial contigs where models trained on KIMI-selected k-mers increased accuracy of viral versus bacterial classification.
Scientific Applications:
- Metagenomic motif discovery: Identification of sequence motifs in metagenomic reads to reveal taxa-specific patterns.
- Microbial community classification: Feature selection for classifiers distinguishing bacteria, viruses, and other microbial groups.
- Virology and microbiology research: Extraction of biologically relevant k-mers to support studies in microbiology, virology, and comparative genomics.
Methodology:
Uses model-X Knockoffs to perform motif discovery and control FDR, selects relevant k-mers from molecular sequences with an adjustable target FDR, and evaluates performance via simulation studies and comparison to the Benjamini–Hochberg and q-value methods.
Topics
Details
- Tool Type:
- library
- Programming Languages:
- C++, R
- Added:
- 1/18/2021
- Last Updated:
- 11/24/2024
Operations
Publications
Bai X, Ren J, Fan Y, Sun F. KIMI: Knockoff Inference for Motif Identification from molecular sequences with controlled false discovery rate. Bioinformatics. 2020;37(6):759-766. doi:10.1093/bioinformatics/btaa912. PMID:33119059. PMCID:PMC8599924.