VariBench
VariBench provides benchmark datasets for development, testing, and performance assessment of computational methods for variation effect prediction.
Key Features:
- Extensive Dataset Collection: Contains 419 distinct datasets from 109 scientific papers comprising over 329 million variants.
- Categorization and Diversity: Datasets are categorized into 20 groups and subgroups covering insertions, deletions, substitutions in coding and non-coding regions, structure-mapped variants, synonymous variants, and benign variants.
- Effect-Specific Datasets: Includes datasets targeting DNA regulatory elements, RNA splicing, and protein properties such as aggregation, binding free energy, disorder, and stability.
- Molecule- and Disease-Specific Datasets: Provides molecule-specific and disease-specific datasets and a dataset for variation phenotype effects.
- Multi-Level Variant Descriptions: Variants are described at DNA, RNA, protein, and protein structural levels.
- Facilitation of Method Development and Testing: Enables training and benchmarking of computational predictors and comparison of method performance against previously published results.
- Performance Comparison Utility: Has been used to compare pathogenicity/tolerance predictors such as PON-P2 with other benchmark studies.
Scientific Applications:
- Predictive model development: Supports development and benchmarking of computational models for variation effect prediction in genomics and bioinformatics.
- Evaluation of pathogenicity/tolerance predictors: Enables evaluation and comparison of pathogenicity/tolerance predictors, exemplified by comparisons involving PON-P2.
- Research on molecular mechanisms: Facilitates studies of DNA regulatory elements, RNA splicing, and protein properties like aggregation and stability.
- Disease and personalized medicine studies: Supports disease-focused analyses and applications in personalized medicine and disease prediction.
- Basic gene function research: Supports basic research on gene function through validated variant datasets.
Methodology:
Datasets are curated from literature, websites, and existing databases and categorized into 20 groups and subgroups, with some redundancy among datasets retained.
Topics
Details
- License:
- Unlicense
- Maturity:
- Mature
- Cost:
- Free of charge
- Tool Type:
- workflow
- Operating Systems:
- Linux, Windows, Mac
- Added:
- 8/9/2019
- Last Updated:
- 6/16/2020
Operations
Publications
Sarkar A, Yang Y, Vihinen M. Variation Benchmark Datasets: Update, Criteria, Quality and Applications. Unknown Journal. 2019. doi:10.1101/634766.
DOI: 10.1101/634766