DIMA
DIMA selects the optimal imputation algorithm for proteomics datasets using data-driven performance evaluations to improve handling of missing values in quantitative proteomics analyses.
Key Features:
- Performance Assessment: DIMA evaluates the performance of multiple imputation algorithms across diverse quantitative proteomics datasets to identify high-performing methods tailored to dataset properties.
- Data-Driven Selection: Using empirical evaluations on 142 quantitative proteomics datasets from the PRoteomics IDEntifications (PRIDE) database and simulated datasets with 5-50% missing values, DIMA determines the most suitable imputation algorithm based on performance metrics.
- Robustness Across MV Types: DIMA handles missing not at random and missing completely at random scenarios to accommodate different missing-value mechanisms.
- Accuracy and Reliability: DIMA consistently identifies an algorithm among the top three performers and achieves a ΔRMSE ≤ 10% in 80% of evaluated cases.
Scientific Applications:
- Proteomics Data Analysis: Optimizing imputation strategies to mitigate the impact of missing values on quantitative proteomics results and downstream analyses.
- Broad Applicability: Applicable to datasets beyond proteomics by adapting to varying proportions and types of missing values.
Methodology:
DIMA evaluates multiple imputation algorithms on real-world (PRIDE) and simulated datasets and recommends the algorithm with the best empirical performance for the specific dataset characteristics.
Topics
Details
- Tool Type:
- library
- Operating Systems:
- Mac, Linux, Windows
- Programming Languages:
- R, MATLAB
- Added:
- 11/3/2021
- Last Updated:
- 11/3/2021
Operations
Publications
Egert J, Brombacher E, Warscheid B, Kreutz C. DIMA: Data-Driven Selection of an Imputation Algorithm. Journal of Proteome Research. 2021;20(7):3489-3496. doi:10.1021/acs.jproteome.1c00119. PMID:34062065.
PMID: 34062065
Funding: - Bundesministerium f?r Bildung und Forschung: FKZ031L0080
- Deutsche Forschungsgemeinschaft: 403222702278002225/SFB 1381, CIBSS-EXC-2189-2100249960-390939984, FOR 2743, INST35/1134-1FUGG, TRR 130
- H2020 Marie Sklodowska-Curie Actions: 812968
- H2020 European Research Council: 648235
Links
Repository
http://github.com/kreutz-lab/OmicsData