GAMUT
GAMUT manages and compares high-throughput Single Nucleotide Polymorphism (SNP) variant data derived from Next Generation Sequencing to enable population-level variant comparison and genotype–phenotype analyses.
Key Features:
- Efficient SNP comparison: Generates probable lists of variants that differentiate two or more populations for applications such as array design, genotype imputation, and cataloging variants in regions of interest.
- Scalable architecture: Implements a client-server architecture with MongoDB as the backend database for scalable variant data management.
- Parallel data uploading: Uses Spark-based parallel data uploading to accelerate ingestion of large genomic datasets.
- Deployment options: Deployable on WildFly application servers and Docker containers for varied computational environments.
- Dynamic querying: Supports queries by free text, chromosome position, and gene name to retrieve subsets of variant data.
- Visualization outputs: Produces bar and pie charts and tabular formats for representing queried variant data.
- Export formats: Provides data export in text, HTML, and JSON formats.
Scientific Applications:
- Population genetics: Comparative analysis of SNP distributions between populations to identify differentiating variants.
- Array design: Selection of candidate SNPs for genotyping array content based on population differentiation.
- Genotype imputation support: Identification and cataloging of variants to inform imputation reference selection and evaluation.
- Region-specific variant cataloging: Aggregation and comparison of variants within genomic regions of interest for functional and association studies.
- Genotype–phenotype analysis: Facilitating exploration of variants potentially underlying phenotypic traits.
Methodology:
Processes SNPs from Next Generation Sequencing; generates probable population-differentiating variant lists; uses a client-server design with MongoDB backend; employs Spark-based parallel data uploading; supports text-, chromosome position-, and gene-name-based querying; produces bar/pie charts and tables and exports in text, HTML, and JSON; deployable on WildFly and Docker.
Topics
Details
- License:
- Not licensed
- Cost:
- Free of charge
- Tool Type:
- library
- Operating Systems:
- Mac, Linux, Windows
- Programming Languages:
- JavaScript, Java
- Added:
- 3/1/2022
- Last Updated:
- 3/1/2022
Operations
Publications
Ramakrishnan EP, Gupta S, Gadhari R, Bharti N, Malviya S, Kasibhatla SM, Kim J, Joshi R. GAMUT: A genomics big data management tool. Journal of Biosciences. 2021;46(4). doi:10.1007/s12038-021-00213-y. PMID:34544908.