UKBCC

UKBCC curates and generates cohorts from UK Biobank (UKB) datasets by mapping and querying heterogeneous encoding schemes across main and associated datasets, including general practitioner clinical data.


Key Features:

  • Python implementation: Implemented in Python to enable programmatic cohort construction and integration with analysis workflows.
  • Search-term based cohort construction: Accepts user-defined search terms to identify participants meeting precise inclusion and exclusion criteria.
  • Multi-dataset filtering: Filters across main UKB datasets and associated datasets such as general practitioner clinical data to combine heterogeneous sources.
  • Module integration: Importable as a module for incorporation into broader analytical pipelines.
  • Automated bulk data handling: Supports automatic downloading and handling of large bulk data files.
  • Replicable cohort creation: Provides sub-modules and functions that enable reproducible cohort generation.

Scientific Applications:

  • Cohort generation: Construct targeted cohorts from UK Biobank for research studies using user-defined search terms.
  • Participant subset identification: Identify participant subsets that satisfy precise selection criteria across encoded, heterogeneous UKB data sources.

Methodology:

Implemented in Python with sub-modules and functions that accept user-defined search terms, perform filtering across main and associated UKB datasets (including general practitioner clinical data), and support automatic downloading of bulk data files.

Topics

Details

License:
Apache-2.0
Programming Languages:
Python
Added:
1/18/2021
Last Updated:
3/6/2021

Operations

Publications

Kiral I, Willems N, Goudey B. UKBCC: a cohort curation package for UK Biobank. Unknown Journal. 2020. doi:10.1101/2020.07.12.199810.

Links