Observational Health Data Sciences and Informatics

Observational Health Data Sciences and Informatics provides a standardized analytics pipeline for development and rigorous validation of prediction models using observational health data.


Key Features:

  • Standardized analytics pipeline: A structured approach for rapid development and validation of prediction models using observational health data.
  • Bias mitigation strategies: Validates phenotypes, specifies target populations, conducts large-scale external validations, and publicly shares analytical source code to reduce sources of bias.
  • Machine learning methods: Implements AdaBoost, random forest, gradient boosting machine, decision tree, L1-regularized logistic regression, and MLP neural network for model development.
  • End-to-end process: Supports steps from problem design through model development and evaluation.
  • Data scale and sources: Applicable to observational datasets such as claims databases (e.g., development on a USA claims database with >20,000 hospitalizations) and supports external validation across multi-country datasets (>45,000 hospitalizations).

Scientific Applications:

  • COVID-19 mortality risk prediction (0 to 30 days): Development and validation of models predicting 0 to 30 day post-hospitalization mortality for COVID-19 patients.
  • Cross-population external validation: External validation of predictive models across datasets from the USA, South Korea, and Spain.
  • Comparative algorithm evaluation: Comparative assessment of AdaBoost, random forest, gradient boosting machine, decision tree, L1-regularized logistic regression, and MLP neural network performance.
  • Model performance assessment: Evaluation of discrimination and calibration during internal and external validations.

Methodology:

Model development used AdaBoost, random forest, gradient boosting machine, decision tree, L1-regularized logistic regression, and MLP neural network; phenotype validation and explicit target population specification were performed; internal and large-scale external validations were conducted using observational claims data (development: >20,000 US hospitalizations; external: >45,000 hospitalizations across South Korea, Spain, and the USA); discrimination and calibration were evaluated and analytical source code was shared.

Topics

Details

Tool Type:
library
Added:
1/18/2021
Last Updated:
1/3/2022

Operations

Publications

Khalid S, Yang C, Blacketer C, Duarte-Salles T, Fernández-Bertolín S, Kim C, Park RW, Park J, Schuemie MJ, Sena AG, Suchard MA, You SC, Rijnbeek PR, Reps JM. A standardized analytics pipeline for reliable and rapid development and validation of prediction models using observational health data. Computer Methods and Programs in Biomedicine. 2021;211:106394. doi:10.1016/j.cmpb.2021.106394. PMID:34560604. PMCID:PMC8420135.

Links