SparkRA
SparkRA accelerates GATK RNA-seq variant calling by distributing the GATK RNA-seq best-practices pipeline with Apache Spark to enable scalable processing of large RNA-seq datasets.
Key Features:
- Apache Spark integration: Implements Apache Spark–based distributed execution of the GATK RNA-seq best-practices variant calling pipeline.
- Scalability: Supports scaling across multiple cores on a single node and across multi-node clusters.
- Single-node performance: On a single node with 20 hyper-threaded cores processing a 32 GB dataset, SparkRA reduced runtime from over five hours to approximately 1 hour 18 minutes (~4× speedup).
- Multi-node performance: On a 16-node cluster (each node with eight single-threaded cores), SparkRA achieved a 7.7× reduction in computation time relative to single-node processing.
- Comparative performance and accuracy: Demonstrated 1.2× faster performance than other contemporary scalable solutions while maintaining equivalent accuracy.
Scientific Applications:
- Genotype–phenotype analysis: Enables large-scale RNA-seq variant analyses to support investigations of genotype–phenotype relationships.
- Variant discovery at scale: Facilitates high-throughput RNA-seq variant calling for large datasets.
- Result corroboration: Supports corroboration of findings from other analytical methods using RNA-seq variant calls.
Methodology:
Implements Apache Spark–based distributed execution of the GATK RNA-seq best-practices variant calling pipeline across single-node multicore and multi-node cluster environments, with performance validated using benchmarks on a 32 GB RNA-seq dataset.
Topics
Details
- Added:
- 1/18/2021
- Last Updated:
- 2/20/2021
Operations
Publications
Al-Ars Z, Wang S, Mushtaq H. SparkRA: Enabling Big Data Scalability for the GATK RNA-seq Pipeline with Apache Spark. Genes. 2020;11(1):53. doi:10.3390/genes11010053. PMID:31947774. PMCID:PMC7016739.