i6mA-stack

i6mA-stack predicts DNA N6-methyladenine (6mA) sites in Rosaceae genomes, specifically Rosa chinensis and Fragaria vesca, to support analysis of epigenetic 6mA distributions.


Key Features:

  • Machine Learning-Based Approach: Employs a two-layer stacking ensemble of machine learning algorithms for 6mA site prediction.
  • Feature Optimization: Uses recursive feature elimination with cross-validation (RFECV) to determine an optimal number of features (ONF) from five DNA sequence encoding schemes: Binary Encoding (BE), Ring-Function-Hydrogen-Chemical Properties (RFHC), Electron-Ion-Interaction Pseudo Potentials of Nucleotides (EIIP), Dinucleotide Physicochemical Properties (DPCP), and Trinucleotide Physicochemical Properties (TPCP).
  • Benchmarking: Demonstrated superior performance versus peer tools in predicting 6mA sites.

Scientific Applications:

  • 6mA site mapping in Rosaceae: Identification of genomic 6mA sites in Rosa chinensis and Fragaria vesca for studies of methylation distribution.
  • Epigenetic function analysis: Support for investigating cellular processes influenced by 6mA, including aspects of epigenetic inheritance.
  • Cross-domain research: Applicability to studies of 6mA roles in both eukaryotic and prokaryotic organisms.

Methodology:

Train a two-layer stacking ensemble using an ONF subset of features selected by RFECV derived from BE, RFHC, EIIP, DPCP, and TPCP encoding schemes.

Topics

Details

Tool Type:
api
Added:
1/18/2021
Last Updated:
2/1/2021

Operations

Publications

Khanal J, Lim DY, Tayara H, Chong KT. i6mA-stack: A stacking ensemble-based computational prediction of DNA N6-methyladenine (6mA) sites in the Rosaceae genome. Genomics. 2021;113(1):582-592. doi:10.1016/j.ygeno.2020.09.054. PMID:33010390.