MLM

MLM improves data splitting for classification models in spectrochemical and biospectroscopic analyses by integrating a random-mutation factor into the Kennard-Stone algorithm to produce more balanced training and test sets.


Key Features:

  • Random-mutation factor: Integrates stochastic mutations into the sample selection process to modify Kennard-Stone outcomes.
  • Extension of Kennard-Stone (KS): Builds upon the KS algorithm for iterative selection of representative samples.
  • Euclidean distance-based selection: Retains KS calculation of Euclidean distances between samples to ensure representative distribution of training and test sets.
  • Balanced training/test sets: Aims to produce more balanced datasets to avoid unbalanced classification metrics.
  • Improved sensitivity and specificity: Reported to improve sensitivity and specificity values relative to random selection (RS) and original KS.
  • Benchmarking and evaluation: Evaluated against RS and KS across simulated scenarios and six real-world biospectroscopic applications.
  • PCA-LDA evaluation: Performance assessed using principal component analysis linear discriminant analysis (PCA-LDA).
  • Application domain: Targeted at spectrochemical analyses and biospectroscopic datasets for classification model development.

Scientific Applications:

  • Spectrochemical analyses: Data splitting for classification models in spectrochemical studies.
  • Biospectroscopic applications: Applied to six real-world biospectroscopic datasets for classifier development and validation.
  • Biomedical classification models: Used to improve sensitivity and specificity in biomedical spectroscopic classification tasks.
  • Simulation studies: Validated across simulated scenarios to benchmark performance against RS and KS.
  • PCA-LDA-based evaluation: Suited for workflows that use principal component analysis linear discriminant analysis for model assessment.

Methodology:

Extends the Kennard-Stone (KS) Euclidean distance-based sample selection by incorporating a random-mutation factor for training/test split selection and assesses performance by comparing RS, KS, and MLM using principal component analysis linear discriminant analysis (PCA-LDA) on simulated scenarios and six real-world biospectroscopic datasets.

Topics

Details

License:
GPL-3.0
Maturity:
Mature
Cost:
Free of charge
Tool Type:
command-line tool
Operating Systems:
Linux, Windows, Mac
Programming Languages:
MATLAB
Added:
8/9/2019
Last Updated:
11/24/2024

Operations

Publications

Morais CLM, Santos MCD, Lima KMG, Martin FL. Improving data splitting for classification applications in spectrochemical analyses employing a random-mutation Kennard-Stone algorithm approach. Bioinformatics. 2019;35(24):5257-5263. doi:10.1093/bioinformatics/btz421. PMID:31116391. PMCID:PMC6954661.

PMID: 31116391
PMCID: PMC6954661
Funding: - Coordenação de Aperfeiçoamento de Pessoal de Nível Superior: 88881.128982/2016-01 - Biotechnology and Biological Sciences Research Council: BB/D010055/1 - Engineering and Physical Sciences Research Council: EP/K023349/1, GR/S75918/01