SEQSpark

SEQSpark leverages Apache Spark to parallelize processing of large-scale sequence-based association studies for rapid detection and analysis of rare variant associations in exome and whole-genome data generated by massively parallel sequencing technologies.


Key Features:

  • Parallel Processing with Apache Spark: Uses Apache Spark's distributed computing to parallelize data quality control, annotation, and association analysis across large datasets.
  • Comprehensive Analysis Capabilities: Implements principal component analysis, single-variant association testing, and rare-variant aggregate association analyses.
  • Scalability and Variant Capacity: Handles datasets with over 9 million variants and whole-genome sequence data.
  • Demonstrated Performance on UK10K: Processed UK10K whole-genome sequence data including data loading, annotation, and multiple association analyses for waist-to-hip ratio in 1.5 hours.
  • Rare Variant Association Detection: Detected a significant association at gene CCDC62 with SKAT-O p = 6.89 × 10^-7, combined multivariate collapsing p = 1.48 × 10^-6, and burden p = 1.48 × 10^-6.
  • Performance Benchmarking: Outperformed Variant Association Tools and PLINK/SEQ in benchmarking, in some cases reducing computation time by up to a hundredth.

Scientific Applications:

  • Large-scale imputation and sequence-based association studies: Applies to large-scale imputation and exome or whole-genome sequence-based association studies aimed at identifying rare susceptibility variants.
  • Genetic epidemiology of complex traits: Supports genetic epidemiology studies of complex diseases and traits, exemplified by analyses of waist-to-hip ratio.
  • Rare variant discovery and aggregate testing: Enables discovery and testing of aggregate rare-variant signals using methods such as SKAT-O, combined multivariate collapsing, and burden tests.

Methodology:

Integrates Apache Spark's distributed computing framework to parallelize data loading, quality control, annotation, and association analyses on large-scale genomic datasets.

Topics

Details

License:
Apache-2.0
Tool Type:
command-line tool
Operating Systems:
Linux, Mac
Programming Languages:
Java, Scala, Perl
Added:
8/1/2018
Last Updated:
1/15/2019

Operations

Data Inputs & Outputs

Publications

Zhang D, Zhao L, Li B, He Z, Wang GT, Liu DJ, Leal SM. SEQSpark: A Complete Analysis Tool for Large-Scale Rare Variant Association Studies Using Whole-Genome and Exome Sequence Data. The American Journal of Human Genetics. 2017;101(1):115-122. doi:10.1016/j.ajhg.2017.05.017. PMID:28669402. PMCID:PMC5501866.

PMID: 28669402
PMCID: PMC5501866
Funding: - National Human Genome Research Institute: HG008972

Documentation

Links