PyGMQL
PyGMQL provides scalable extraction, manipulation, and analysis of region-based genomic datasets and their associated metadata using the GMQL (Genomics Massive Query Language) big data management system to support large-scale genomic studies.
Key Features:
- Scalability: Handles region-based genomic files and their associated metadata across arbitrary clusters, enabling processing of millions of genomic regions.
- Data Interoperability: Integrates GMQL set-oriented query semantics with Python through Apache Spark to enable set-oriented operations within Python workflows.
- Distribution Transparency and Query Outsourcing: Allows addressing remote datasets as if local and supports outsourcing queries to remote or cloud-based GMQL engine installations.
- Expressive Functions for Data Manipulation: Provides a comprehensive set of functions for manipulation of genomic regions and their metadata to support complex analyses.
- Reproducibility and Scalability: Supports reproducible, scalable tertiary data extraction and analysis pipelines demonstrated across biological scenarios.
Scientific Applications:
- Heterogeneous processed data analysis: Enables analysis of heterogeneous processed genomic datasets to derive biological and clinical insights.
- Genome-wide association studies (GWAS): Supports high-throughput region-based processing required for GWAS.
- Comparative genomics: Facilitates comparative genomics analyses across large cohorts of experiments.
- Personalized medicine: Enables processing workflows applicable to personalized medicine studies.
Methodology:
Leverages the Apache Spark engine underlying GMQL for distributed computing, integrates GMQL query execution with Python, and supports distribution transparency and outsourcing of queries to remote or cloud-based GMQL engines.
Topics
Details
- License:
- Apache-2.0
- Programming Languages:
- Python
- Added:
- 1/14/2020
- Last Updated:
- 12/11/2020
Operations
Publications
Nanni L, Pinoli P, Canakoglu A, Ceri S. PyGMQL: scalable data extraction and analysis for heterogeneous genomic datasets. BMC Bioinformatics. 2019;20(1). doi:10.1186/s12859-019-3159-9. PMID:31703553.
Downloads
- Container filehttps://hub.docker.com/r/gecopolimi/pygmql