VCaScale

VCaScale performs scalable pre-processing and variant calling of next-generation sequencing (NGS) data to enable deployment of DeepVariant and other deep learning–based variant callers across CPU-only and hybrid CPU+GPU clusters.


Key Features:

  • Scalability and Efficiency: Integrates tightly coupled pre-processing stages to optimize performance for large-scale genomic studies on CPU-only and hybrid CPU+GPU clusters.
  • Apache Spark Integration: Uses Apache Spark built-in functions to sort reads by coordinates and mark duplicates while exploiting Spark in-memory processing to reduce computational overhead.
  • Native Spark-based Workflow: Employs a native Spark-based workflow with Python and Apache Arrow for columnar in-memory data transformations, improving pre-processing performance by more than two-fold versus state-of-the-art implementations.
  • High Performance on HPC Clusters: Optimized for high-performance computing (HPC) clusters and supports standardized Apache Arrow data representations to facilitate data transfer between workflow stages.
  • Deep learning-based Variant Calling Support: Leverages DeepVariant and other deep learning methods that surpass GATK HaplotypeCaller, Strelka2, and Freebayes in accuracy while accommodating their higher computational demands.

Scientific Applications:

  • Large-scale cohort variant calling: Enables scalable pre-processing and variant calling for large NGS cohorts.
  • High-accuracy variant detection: Facilitates application of DeepVariant and other deep learning callers for improved variant detection accuracy.
  • Basic research and clinical genomics: Supports projects that require precise and scalable variant analysis for research and clinical applications.

Methodology:

Tightly integrated workflow that couples data pre-processing and variant calling; uses Apache Spark built-in functions to sort reads by coordinates and mark duplicates, Python for programmability, and Apache Arrow for columnar in-memory data transformations to minimize I/O bottlenecks and maximize computational efficiency while targeting CPU-only and hybrid CPU+GPU clusters and leveraging DeepVariant for variant calling.

Topics

Details

License:
Apache-2.0
Cost:
Free of charge
Tool Type:
command-line tool
Operating Systems:
Mac, Windows, Linux
Programming Languages:
Python, Shell
Added:
1/21/2022
Last Updated:
1/21/2022

Operations

Publications

Ahmad T, Al Ars Z, Hofstee HP. VC@Scale: Scalable and high-performance variant calling on cluster environments. GigaScience. 2021;10(9). doi:10.1093/gigascience/giab057. PMID:34494101. PMCID:PMC8424057.