R.ROSETTA
R.ROSETTA implements interpretable rule-based classification and statistical characterization of models for bioinformatics datasets using rough set theory.
Key Features:
- Rule-Based Classification Models: Builds interpretable, non-linear classifiers by inducing rules from data using rough set theory.
- Rough Set Theory: Applies rough set theory to handle vagueness and uncertainty in decision-table formatted data.
- Combinatorial Statistics: Computes combinatorial statistics for rules to quantify support and statistical characteristics of model components.
- Data Balancing (Undersampling): Performs data balancing through undersampling to address class imbalance.
- Rule P-value Estimation: Estimates P values for individual rules to assess statistical significance.
- Support Set Retrieval: Retrieves support sets associated with induced rules.
- Model Merging: Merges rule-based models to consolidate rule sets across analyses.
- Prediction of External Classes: Predicts class labels for external or new samples using induced rule sets.
- Visualization Tools: Provides visualization routines to inspect rules and model structure.
- Synthetic Data Generation: Generates synthetic data for analysis and testing.
- Statistical Properties Reporting: Reports statistical properties for models and their components to aid interpretation.
Scientific Applications:
- Life sciences interpretability: Applied where interpretable machine learning is required to explain predictive mechanisms in biological data.
- Transcriptome analysis (autism case-control): Used on transcriptome datasets from autism case-control studies to generate hypotheses on co-predictive feature mechanisms that distinguish phenotype classes and to identify neurodevelopmental and autism-related genes.
Methodology:
The methodology uses rough set theory and rule-based modeling on decision-table formatted data, with explicit computational procedures including combinatorial statistics for rule evaluation, undersampling for data balancing, estimation of rule P values, retrieval of rule support sets, model merging, prediction of external classes, visualization routines, and synthetic data generation.
Topics
Details
- Tool Type:
- library
- Programming Languages:
- R
- Added:
- 11/29/2021
- Last Updated:
- 11/29/2021
Operations
Publications
Garbulowski M, Diamanti K, Smolińska K, Baltzer N, Stoll P, Bornelöv S, Øhrn A, Feuk L, Komorowski J. R.ROSETTA: an interpretable machine learning framework. BMC Bioinformatics. 2021;22(1). doi:10.1186/s12859-021-04049-z. PMID:33676405. PMCID:PMC7937228.