SparkRA

SparkRA accelerates GATK RNA-seq variant calling by distributing the GATK RNA-seq best-practices pipeline with Apache Spark to enable scalable processing of large RNA-seq datasets.


Key Features:

  • Apache Spark integration: Implements Apache Spark–based distributed execution of the GATK RNA-seq best-practices variant calling pipeline.
  • Scalability: Supports scaling across multiple cores on a single node and across multi-node clusters.
  • Single-node performance: On a single node with 20 hyper-threaded cores processing a 32 GB dataset, SparkRA reduced runtime from over five hours to approximately 1 hour 18 minutes (~4× speedup).
  • Multi-node performance: On a 16-node cluster (each node with eight single-threaded cores), SparkRA achieved a 7.7× reduction in computation time relative to single-node processing.
  • Comparative performance and accuracy: Demonstrated 1.2× faster performance than other contemporary scalable solutions while maintaining equivalent accuracy.

Scientific Applications:

  • Genotype–phenotype analysis: Enables large-scale RNA-seq variant analyses to support investigations of genotype–phenotype relationships.
  • Variant discovery at scale: Facilitates high-throughput RNA-seq variant calling for large datasets.
  • Result corroboration: Supports corroboration of findings from other analytical methods using RNA-seq variant calls.

Methodology:

Implements Apache Spark–based distributed execution of the GATK RNA-seq best-practices variant calling pipeline across single-node multicore and multi-node cluster environments, with performance validated using benchmarks on a 32 GB RNA-seq dataset.

Topics

Details

Added:
1/18/2021
Last Updated:
2/20/2021

Operations

Publications

Al-Ars Z, Wang S, Mushtaq H. SparkRA: Enabling Big Data Scalability for the GATK RNA-seq Pipeline with Apache Spark. Genes. 2020;11(1):53. doi:10.3390/genes11010053. PMID:31947774. PMCID:PMC7016739.