Observational Health Data Sciences and Informatics
Observational Health Data Sciences and Informatics provides a standardized analytics pipeline for development and rigorous validation of prediction models using observational health data.
Key Features:
- Standardized analytics pipeline: A structured approach for rapid development and validation of prediction models using observational health data.
- Bias mitigation strategies: Validates phenotypes, specifies target populations, conducts large-scale external validations, and publicly shares analytical source code to reduce sources of bias.
- Machine learning methods: Implements AdaBoost, random forest, gradient boosting machine, decision tree, L1-regularized logistic regression, and MLP neural network for model development.
- End-to-end process: Supports steps from problem design through model development and evaluation.
- Data scale and sources: Applicable to observational datasets such as claims databases (e.g., development on a USA claims database with >20,000 hospitalizations) and supports external validation across multi-country datasets (>45,000 hospitalizations).
Scientific Applications:
- COVID-19 mortality risk prediction (0 to 30 days): Development and validation of models predicting 0 to 30 day post-hospitalization mortality for COVID-19 patients.
- Cross-population external validation: External validation of predictive models across datasets from the USA, South Korea, and Spain.
- Comparative algorithm evaluation: Comparative assessment of AdaBoost, random forest, gradient boosting machine, decision tree, L1-regularized logistic regression, and MLP neural network performance.
- Model performance assessment: Evaluation of discrimination and calibration during internal and external validations.
Methodology:
Model development used AdaBoost, random forest, gradient boosting machine, decision tree, L1-regularized logistic regression, and MLP neural network; phenotype validation and explicit target population specification were performed; internal and large-scale external validations were conducted using observational claims data (development: >20,000 US hospitalizations; external: >45,000 hospitalizations across South Korea, Spain, and the USA); discrimination and calibration were evaluated and analytical source code was shared.
Topics
Details
- Tool Type:
- library
- Added:
- 1/18/2021
- Last Updated:
- 1/3/2022
Operations
Publications
Khalid S, Yang C, Blacketer C, Duarte-Salles T, Fernández-Bertolín S, Kim C, Park RW, Park J, Schuemie MJ, Sena AG, Suchard MA, You SC, Rijnbeek PR, Reps JM. A standardized analytics pipeline for reliable and rapid development and validation of prediction models using observational health data. Computer Methods and Programs in Biomedicine. 2021;211:106394. doi:10.1016/j.cmpb.2021.106394. PMID:34560604. PMCID:PMC8420135.