MLM
MLM improves data splitting for classification models in spectrochemical and biospectroscopic analyses by integrating a random-mutation factor into the Kennard-Stone algorithm to produce more balanced training and test sets.
Key Features:
- Random-mutation factor: Integrates stochastic mutations into the sample selection process to modify Kennard-Stone outcomes.
- Extension of Kennard-Stone (KS): Builds upon the KS algorithm for iterative selection of representative samples.
- Euclidean distance-based selection: Retains KS calculation of Euclidean distances between samples to ensure representative distribution of training and test sets.
- Balanced training/test sets: Aims to produce more balanced datasets to avoid unbalanced classification metrics.
- Improved sensitivity and specificity: Reported to improve sensitivity and specificity values relative to random selection (RS) and original KS.
- Benchmarking and evaluation: Evaluated against RS and KS across simulated scenarios and six real-world biospectroscopic applications.
- PCA-LDA evaluation: Performance assessed using principal component analysis linear discriminant analysis (PCA-LDA).
- Application domain: Targeted at spectrochemical analyses and biospectroscopic datasets for classification model development.
Scientific Applications:
- Spectrochemical analyses: Data splitting for classification models in spectrochemical studies.
- Biospectroscopic applications: Applied to six real-world biospectroscopic datasets for classifier development and validation.
- Biomedical classification models: Used to improve sensitivity and specificity in biomedical spectroscopic classification tasks.
- Simulation studies: Validated across simulated scenarios to benchmark performance against RS and KS.
- PCA-LDA-based evaluation: Suited for workflows that use principal component analysis linear discriminant analysis for model assessment.
Methodology:
Extends the Kennard-Stone (KS) Euclidean distance-based sample selection by incorporating a random-mutation factor for training/test split selection and assesses performance by comparing RS, KS, and MLM using principal component analysis linear discriminant analysis (PCA-LDA) on simulated scenarios and six real-world biospectroscopic datasets.
Topics
Details
- License:
- GPL-3.0
- Maturity:
- Mature
- Cost:
- Free of charge
- Tool Type:
- command-line tool
- Operating Systems:
- Linux, Windows, Mac
- Programming Languages:
- MATLAB
- Added:
- 8/9/2019
- Last Updated:
- 11/24/2024
Operations
Publications
Morais CLM, Santos MCD, Lima KMG, Martin FL. Improving data splitting for classification applications in spectrochemical analyses employing a random-mutation Kennard-Stone algorithm approach. Bioinformatics. 2019;35(24):5257-5263. doi:10.1093/bioinformatics/btz421. PMID:31116391. PMCID:PMC6954661.