Nm-Nano

Nm-Nano predicts 2′-O-methylation (Nm) sites in RNA from Oxford Nanopore direct RNA sequencing data to enable mapping of Nm modifications in human cell lines.


Key Features:

  • Machine Learning Models: Integrates supervised Extreme Gradient Boosting (XGBoost) and Random Forest (RF) models to classify Nm-modified versus unmodified sites.
  • k-mer Embedding: Employs k-mer embedding techniques, with RF incorporating dense vector representations of RNA k-mers generated by the word2vec technique to capture sequence-specific features.
  • Feature Extraction: Trains models on features derived from modified and unmodified nanopore signals together with their corresponding k-mers obtained through base-calling.
  • Performance and Validation: Reported accuracies of 99% for XGBoost and 92% for RF using integrated validation on Hela and Hek293 benchmark datasets with a 50% train / 50% test split.
  • Biological Insight Capability: Enables identification of frequently Nm-modified genes and supports downstream functional enrichment analyses linking Nm to immune response, C3HC4-type RING finger domain binding, antigen processing and presentation (class I MHC), glycolysis/gluconeogenesis, and protein localization.

Scientific Applications:

  • Mapping Nm in human cell lines: Detects and maps Nm sites in Oxford Nanopore direct RNA-seq data from human cell lines, including Hela and Hek293.
  • Gene-level modification profiling: Identified 125 frequently Nm-modified genes in Hela and 61 top Nm-modified genes in Hek293 for downstream analysis.
  • Functional and molecular studies: Facilitates investigation of Nm roles in tRNA functionality, mRNA protection against degradation by DXO, and rRNA biogenesis and specificity.

Methodology:

Features are extracted from modified and unmodified nanopore signals and base-called k-mers; RF uses word2vec-generated dense k-mer vectors and XGBoost uses k-mer-based features; both supervised models (XGBoost and RF) are trained and evaluated on Hela and Hek293 benchmark datasets with a 50% train / 50% test split, yielding reported accuracies of 99% (XGBoost) and 92% (RF).

Topics

Details

License:
MIT
Cost:
Free of charge
Tool Type:
command-line tool
Operating Systems:
Mac, Linux, Windows
Programming Languages:
Python
Added:
9/17/2022
Last Updated:
11/24/2024

Operations

Data Inputs & Outputs

Base-calling

Outputs

    Publications

    Salem DH, Ariyur A, Daulatabad SV, Mir Q, Janga SC. Nm-Nano: A Machine Learning Framework for Transcriptome-Wide Single Molecule Mapping of 2´-O-Methylation (Nm) Sites in Nanopore Direct RNA Sequencing Datasets. Unknown Journal. 2022. doi:10.1101/2022.01.03.473214.