R.ROSETTA

R.ROSETTA implements interpretable rule-based classification and statistical characterization of models for bioinformatics datasets using rough set theory.


Key Features:

  • Rule-Based Classification Models: Builds interpretable, non-linear classifiers by inducing rules from data using rough set theory.
  • Rough Set Theory: Applies rough set theory to handle vagueness and uncertainty in decision-table formatted data.
  • Combinatorial Statistics: Computes combinatorial statistics for rules to quantify support and statistical characteristics of model components.
  • Data Balancing (Undersampling): Performs data balancing through undersampling to address class imbalance.
  • Rule P-value Estimation: Estimates P values for individual rules to assess statistical significance.
  • Support Set Retrieval: Retrieves support sets associated with induced rules.
  • Model Merging: Merges rule-based models to consolidate rule sets across analyses.
  • Prediction of External Classes: Predicts class labels for external or new samples using induced rule sets.
  • Visualization Tools: Provides visualization routines to inspect rules and model structure.
  • Synthetic Data Generation: Generates synthetic data for analysis and testing.
  • Statistical Properties Reporting: Reports statistical properties for models and their components to aid interpretation.

Scientific Applications:

  • Life sciences interpretability: Applied where interpretable machine learning is required to explain predictive mechanisms in biological data.
  • Transcriptome analysis (autism case-control): Used on transcriptome datasets from autism case-control studies to generate hypotheses on co-predictive feature mechanisms that distinguish phenotype classes and to identify neurodevelopmental and autism-related genes.

Methodology:

The methodology uses rough set theory and rule-based modeling on decision-table formatted data, with explicit computational procedures including combinatorial statistics for rule evaluation, undersampling for data balancing, estimation of rule P values, retrieval of rule support sets, model merging, prediction of external classes, visualization routines, and synthetic data generation.

Topics

Details

Tool Type:
library
Programming Languages:
R
Added:
11/29/2021
Last Updated:
11/29/2021

Operations

Publications

Garbulowski M, Diamanti K, Smolińska K, Baltzer N, Stoll P, Bornelöv S, Øhrn A, Feuk L, Komorowski J. R.ROSETTA: an interpretable machine learning framework. BMC Bioinformatics. 2021;22(1). doi:10.1186/s12859-021-04049-z. PMID:33676405. PMCID:PMC7937228.

PMID: 33676405
PMCID: PMC7937228
Funding: - Foundation for the National Institutes of Health: NO 0925-0001 - Vetenskapsrådet: 2017-01861

Links