GMQL
GMQL enables advanced querying and integrative analysis of Next Generation Sequencing (NGS)-derived genomic, transcriptomic, and epigenomic region data by linking genomic features to experimental, biological, and clinical metadata for large-scale tertiary analyses.
Key Features:
- Declarative query language: A high-level language that abstracts genomic region data together with associated experimental, biological, and clinical metadata.
- Genomic Data Model (GDM): A unified framework that links genomic feature data with relevant metadata and describes semantically heterogeneous (epi)genomic datasets.
- File format support: Native interoperability with SAM, VCF, NARROWPEAK, and BED formats.
- Interoperability: Integration of diverse genomic, transcriptomic, and epigenomic datasets for cross-dataset queries.
- Scalability and execution engine: Implementation on the Hadoop framework and Apache Pig to enable scalable processing of large datasets.
- Intermediate representation: An intermediate representation that supports alternative implementations including Spark, Flink, and SciDB.
- Repository abstraction: A technology-independent abstraction capable of interfacing with local file systems, Hadoop File Systems, and databases.
- Downstream operation: Designed to operate downstream of raw data preprocessing pipelines and to query across thousands of heterogeneous datasets and samples.
- Tertiary data analysis: Support for tertiary analyses that investigate interactions and cooperations among different (epi)genomic regions and their products.
- Domain-specific computations: Enables development and execution of domain-specific computations and knowledge discovery workflows.
Scientific Applications:
- Integrative analysis of ENCODE data: Cross-dataset queries and integration of ENCODE genomic and epigenomic datasets.
- Roadmap Epigenomics integration: Comparative and integrative analyses of Roadmap Epigenomics datasets across tissues and marks.
- TCGA data analysis: Integration of TCGA genomic and transcriptomic data for tertiary-level studies of cancer genomics.
- Large-scale NGS studies: Big-data genomic analyses spanning thousands of heterogeneous samples for regulatory interaction and epigenomic profiling.
Methodology:
GMQL applies a high-level declarative query language over the Genomic Data Model (GDM), is implemented on Hadoop and Apache Pig, employs an intermediate representation enabling Spark, Flink, and SciDB backends, and uses a technology-independent repository abstraction for local file systems, Hadoop File Systems, and databases while operating downstream of raw data preprocessing pipelines.
Topics
Details
- Tool Type:
- command-line tool
- Operating Systems:
- Linux
- Programming Languages:
- Racket
- Added:
- 8/3/2017
- Last Updated:
- 11/24/2024
Operations
Publications
Masseroli M, Pinoli P, Venco F, Kaitoua A, Jalili V, Palluzzi F, Muller H, Ceri S. GenoMetric Query Language: a novel approach to large-scale genomic data management. Bioinformatics. 2015;31(12):1881-1888. doi:10.1093/bioinformatics/btv048. PMID:25649616.
Masseroli M, Canakoglu A, Pinoli P, Kaitoua A, Gulino A, Horlova O, Nanni L, Bernasconi A, Perna S, Stamoulakatou E, Ceri S. Processing of big heterogeneous genomic datasets for tertiary analysis of Next Generation Sequencing data. Bioinformatics. 2018;35(5):729-736. doi:10.1093/bioinformatics/bty688. PMID:30101316.
Masseroli M, Kaitoua A, Pinoli P, Ceri S. Modeling and interoperability of heterogeneous genomic big data for integrative processing and querying. Methods. 2016;111:3-11. doi:10.1016/j.ymeth.2016.09.002. PMID:27637471.