SEQSpark
SEQSpark leverages Apache Spark to parallelize processing of large-scale sequence-based association studies for rapid detection and analysis of rare variant associations in exome and whole-genome data generated by massively parallel sequencing technologies.
Key Features:
- Parallel Processing with Apache Spark: Uses Apache Spark's distributed computing to parallelize data quality control, annotation, and association analysis across large datasets.
- Comprehensive Analysis Capabilities: Implements principal component analysis, single-variant association testing, and rare-variant aggregate association analyses.
- Scalability and Variant Capacity: Handles datasets with over 9 million variants and whole-genome sequence data.
- Demonstrated Performance on UK10K: Processed UK10K whole-genome sequence data including data loading, annotation, and multiple association analyses for waist-to-hip ratio in 1.5 hours.
- Rare Variant Association Detection: Detected a significant association at gene CCDC62 with SKAT-O p = 6.89 × 10^-7, combined multivariate collapsing p = 1.48 × 10^-6, and burden p = 1.48 × 10^-6.
- Performance Benchmarking: Outperformed Variant Association Tools and PLINK/SEQ in benchmarking, in some cases reducing computation time by up to a hundredth.
Scientific Applications:
- Large-scale imputation and sequence-based association studies: Applies to large-scale imputation and exome or whole-genome sequence-based association studies aimed at identifying rare susceptibility variants.
- Genetic epidemiology of complex traits: Supports genetic epidemiology studies of complex diseases and traits, exemplified by analyses of waist-to-hip ratio.
- Rare variant discovery and aggregate testing: Enables discovery and testing of aggregate rare-variant signals using methods such as SKAT-O, combined multivariate collapsing, and burden tests.
Methodology:
Integrates Apache Spark's distributed computing framework to parallelize data loading, quality control, annotation, and association analyses on large-scale genomic datasets.
Topics
Details
- License:
- Apache-2.0
- Tool Type:
- command-line tool
- Operating Systems:
- Linux, Mac
- Programming Languages:
- Java, Scala, Perl
- Added:
- 8/1/2018
- Last Updated:
- 1/15/2019
Operations
Data Inputs & Outputs
Genetic variation analysis
Publications
Zhang D, Zhao L, Li B, He Z, Wang GT, Liu DJ, Leal SM. SEQSpark: A Complete Analysis Tool for Large-Scale Rare Variant Association Studies Using Whole-Genome and Exome Sequence Data. The American Journal of Human Genetics. 2017;101(1):115-122. doi:10.1016/j.ajhg.2017.05.017. PMID:28669402. PMCID:PMC5501866.