VariBench

VariBench provides benchmark datasets for development, testing, and performance assessment of computational methods for variation effect prediction.


Key Features:

  • Extensive Dataset Collection: Contains 419 distinct datasets from 109 scientific papers comprising over 329 million variants.
  • Categorization and Diversity: Datasets are categorized into 20 groups and subgroups covering insertions, deletions, substitutions in coding and non-coding regions, structure-mapped variants, synonymous variants, and benign variants.
  • Effect-Specific Datasets: Includes datasets targeting DNA regulatory elements, RNA splicing, and protein properties such as aggregation, binding free energy, disorder, and stability.
  • Molecule- and Disease-Specific Datasets: Provides molecule-specific and disease-specific datasets and a dataset for variation phenotype effects.
  • Multi-Level Variant Descriptions: Variants are described at DNA, RNA, protein, and protein structural levels.
  • Facilitation of Method Development and Testing: Enables training and benchmarking of computational predictors and comparison of method performance against previously published results.
  • Performance Comparison Utility: Has been used to compare pathogenicity/tolerance predictors such as PON-P2 with other benchmark studies.

Scientific Applications:

  • Predictive model development: Supports development and benchmarking of computational models for variation effect prediction in genomics and bioinformatics.
  • Evaluation of pathogenicity/tolerance predictors: Enables evaluation and comparison of pathogenicity/tolerance predictors, exemplified by comparisons involving PON-P2.
  • Research on molecular mechanisms: Facilitates studies of DNA regulatory elements, RNA splicing, and protein properties like aggregation and stability.
  • Disease and personalized medicine studies: Supports disease-focused analyses and applications in personalized medicine and disease prediction.
  • Basic gene function research: Supports basic research on gene function through validated variant datasets.

Methodology:

Datasets are curated from literature, websites, and existing databases and categorized into 20 groups and subgroups, with some redundancy among datasets retained.

Topics

Details

License:
Unlicense
Maturity:
Mature
Cost:
Free of charge
Tool Type:
workflow
Operating Systems:
Linux, Windows, Mac
Added:
8/9/2019
Last Updated:
6/16/2020

Operations

Publications

Sarkar A, Yang Y, Vihinen M. Variation Benchmark Datasets: Update, Criteria, Quality and Applications. Unknown Journal. 2019. doi:10.1101/634766.