PyGMQL

PyGMQL provides scalable extraction, manipulation, and analysis of region-based genomic datasets and their associated metadata using the GMQL (Genomics Massive Query Language) big data management system to support large-scale genomic studies.


Key Features:

  • Scalability: Handles region-based genomic files and their associated metadata across arbitrary clusters, enabling processing of millions of genomic regions.
  • Data Interoperability: Integrates GMQL set-oriented query semantics with Python through Apache Spark to enable set-oriented operations within Python workflows.
  • Distribution Transparency and Query Outsourcing: Allows addressing remote datasets as if local and supports outsourcing queries to remote or cloud-based GMQL engine installations.
  • Expressive Functions for Data Manipulation: Provides a comprehensive set of functions for manipulation of genomic regions and their metadata to support complex analyses.
  • Reproducibility and Scalability: Supports reproducible, scalable tertiary data extraction and analysis pipelines demonstrated across biological scenarios.

Scientific Applications:

  • Heterogeneous processed data analysis: Enables analysis of heterogeneous processed genomic datasets to derive biological and clinical insights.
  • Genome-wide association studies (GWAS): Supports high-throughput region-based processing required for GWAS.
  • Comparative genomics: Facilitates comparative genomics analyses across large cohorts of experiments.
  • Personalized medicine: Enables processing workflows applicable to personalized medicine studies.

Methodology:

Leverages the Apache Spark engine underlying GMQL for distributed computing, integrates GMQL query execution with Python, and supports distribution transparency and outsourcing of queries to remote or cloud-based GMQL engines.

Topics

Details

License:
Apache-2.0
Programming Languages:
Python
Added:
1/14/2020
Last Updated:
12/11/2020

Operations

Publications

Nanni L, Pinoli P, Canakoglu A, Ceri S. PyGMQL: scalable data extraction and analysis for heterogeneous genomic datasets. BMC Bioinformatics. 2019;20(1). doi:10.1186/s12859-019-3159-9. PMID:31703553.

PMID: 31703553
PMCID: PMC6842186
Funding: - H2020 European Research Council: 693174

Downloads