iProbiotics

iProbiotics predicts probiotic properties from whole-genome primary sequences using k-mer compositional features and machine learning to identify genomic determinants of probiotic function.


Key Features:

  • Comprehensive dataset utilization: Uses genomic data compiled from the Probiotic Database (PROBIO) and extensive literature surveys.
  • K-mer compositional analysis: Computes k-mer frequencies on strain genomes for k = 2–8 nucleotides to identify oligonucleotide compositions enriched in probiotic genomes.
  • Incremental Feature Selection (IFS): Refines an initial set of 87,376 k-mers using IFS to isolate 184 core features that maximize model performance.
  • High prediction accuracy: Reports predictive performance with accuracy 97.77% and area under the curve (AUC) 98.00%.
  • Functional genomic analysis: Integrates annotations from Gene Ontology (GO), Kyoto Encyclopedia of Genes and Genomes (KEGG), and RAST and examines genes linked to host gastrointestinal survival/settlement, carbohydrate utilization, drug resistance, and virulence factors.
  • Insight into probiotic mechanisms: Indicates probiotic function is associated with combinations of k-mer genomic components rather than single-gene determinants.

Scientific Applications:

  • Probiotic strain identification: Identifies candidate probiotic strains, including lactic acid bacteria consortia commonly found in food, from whole-genome sequences.
  • Experimental prioritization: Prioritizes strains for experimental validation by predicting probiotic properties from genomic features.
  • Functional interpretation: Links predictive k-mer features to GO, KEGG, and RAST annotations and to genes involved in gastrointestinal survival, carbohydrate utilization, drug resistance, and virulence.
  • Microbiome research and probiotic development: Supports genomic analyses that inform microbiome studies and probiotic development efforts via high-accuracy predictions.

Methodology:

Selects a comprehensive dataset (PROBIO and literature), performs k-mer analysis (k = 2–8), applies Incremental Feature Selection to reduce 87,376 k-mers to 184 core features, and integrates GO, KEGG, and RAST annotations for functional analysis.

Topics

Details

Cost:
Free of charge
Tool Type:
web application
Operating Systems:
Mac, Linux, Windows
Added:
5/24/2022
Last Updated:
5/24/2022

Operations

Publications

Sun Y, Li H, Zheng L, Li J, Hong Y, Liang P, Kwok L, Zuo Y, Zhang W, Zhang H. iProbiotics: a machine learning platform for rapid identification of probiotic properties from whole-genome primary sequences. Briefings in Bioinformatics. 2021;23(1). doi:10.1093/bib/bbab477. PMID:34849572.

PMID: 34849572
Funding: - National Natural Science Foundation of China: 31922071, 62061034 - National Natural Science Foundation of Inner Mongolia: 2021ZD08 - Program for Young Talents of Science and Technology in Universities of Inner Mongolia Autonomous Region: NJYT-18-B01 - State Key Laboratory of Reproductive Regulation and Breeding of Grassland Livestock: 2019ZD031

Downloads