MOLGENIS

MOLGENIS enables integration, harmonization, and model-driven generation of biomedical data management and analysis infrastructures to support pooled analyses of omics and clinical datasets.


Key Features:

  • Data Integration and Pooling: MOLGENIS/connect semi-automatically pools data from multiple sources such as biobanks and patient registries, using ontology-based query expansion to identify relevant source attributes.
  • Algorithmic Mapping and Transformation: The system generates algorithms that transform source attributes into a common target DataSchema, handling unit conversion, categorical value matching, and complex calculations such as BMI.
  • Model-Driven Development: A simple modeling language defines biological data structures which the generator suite translates into databases, exchange formats, and scriptable interfaces.
  • Customization and Scalability: A plug-in mechanism allows customization of the generator suite and generated products and supports rapid prototyping for next-generation sequencing, GWAS, QTL analysis, proteomics, and biobanking applications.
  • Development Efficiency: Models typically of around 500 lines of XML replace thousands of lines of hand-written code, enabling reuse of modeling languages and generators across applications.
  • Optimized Back-end and Interfaces: Each MOLGENIS application includes an optimized database back-end and provides programming interfaces in R and Java, with SOAP, REST/JSON, and RDF endpoints and a tab-delimited file format for data exchange.

Scientific Applications:

  • Data pooling and harmonization: Pooling and harmonization of cohort, biobank, and patient registry data to increase statistical power in multi-center and retrospective studies.
  • Genomics and omics analyses: Support for next-generation sequencing, GWAS, QTL analysis, and proteomics data management and prototyping.
  • Biobanking and cohort management: Integration and standardized exchange of biobanking datasets and associated clinical metadata.
  • Automated data transformation: Generation of attribute-mapping algorithms and standardized exchange formats for downstream bioinformatics workflows.

Methodology:

MOLGENIS/connect performs semi-automatic pooling using ontology-based query expansion; it generates attribute-mapping algorithms that perform unit conversions, categorical value matching, and calculations (e.g., BMI) to map to a target DataSchema; a model-driven approach uses a simple modeling language and a generator suite to produce databases, exchange formats, and scriptable interfaces, with customization via plug-ins and programmatic access through R, Java, SOAP, REST/JSON, RDF and a tab-delimited file format.

Topics

Collections

Details

License:
LGPL-3.0
Maturity:
Mature
Cost:
Free of charge
Tool Type:
web application
Operating Systems:
Linux, Windows, Mac
Programming Languages:
R, Java, SQL
Added:
12/2/2015
Last Updated:
11/24/2024

Operations

Publications

Pang C, van Enckevort D, de Haan M, Kelpin F, Jetten J, Hendriksen D, de Boer T, Charbon B, Winder E, van der Velde KJ, Doiron D, Fortier I, Hillege H, Swertz MA. MOLGENIS/connect: a system for semi-automatic integration of heterogeneous phenotype data with applications in biobanks. Bioinformatics. 2016;32(14):2176-2183. doi:10.1093/bioinformatics/btw155. PMID:27153686. PMCID:PMC4937195.

Swertz MA, Dijkstra M, Adamusiak T, van der Velde JK, Kanterakis A, Roos ET, Lops J, Thorisson GA, Arends D, Byelas G, Muilu J, Brookes AJ, de Brock EO, Jansen RC, Parkinson H. The MOLGENIS toolkit: rapid prototyping of biosoftware at the push of a button. BMC Bioinformatics. 2010;11(S12). doi:10.1186/1471-2105-11-s12-s12. PMID:21210979. PMCID:PMC3040526.

Documentation

Downloads