Genomic benchmarks
Genomic benchmarks provides a curated collection of genomic sequence datasets and baseline deep learning resources for benchmarking classification of regulatory elements such as promoters, enhancers, and open chromatin regions across human, mouse, and roundworm.
Key Features:
- Curated Datasets: Nine curated datasets focused on regulatory elements (promoters, enhancers, and open chromatin regions) from human, mouse, and roundworm.
- Baseline Model: A simple convolutional neural network (CNN) is provided as a baseline model for performance comparison.
- Integration with Deep Learning Libraries: An interface compatible with commonly used deep learning libraries is provided to enable integration into existing workflows.
- Training Framework: A training framework is implemented to support reproducible and comparable machine learning experiments.
Scientific Applications:
- Genome Annotation: Use curated datasets to improve annotation of genomic regions with regulatory function.
- Identification of Functional Elements: Benchmark classification methods for detecting promoters, enhancers, and open chromatin regions.
- Machine Learning Research in Genomics: Provide standardized datasets and a baseline model to enable comparative evaluation of deep learning methods.
Methodology:
Datasets were constructed by mining publicly available databases and integrating existing data from published articles.
Topics
Details
- License:
- Apache-2.0
- Cost:
- Free of charge
- Tool Type:
- library
- Operating Systems:
- Mac, Linux, Windows
- Programming Languages:
- Python
- Added:
- 1/10/2024
- Last Updated:
- 11/24/2024
Operations
Publications
Grešová K, Martinek V, Čechák D, Šimeček P, Alexiou P. Genomic benchmarks: a collection of datasets for genomic sequence classification. BMC Genomic Data. 2023;24(1). doi:10.1186/s12863-023-01123-8. PMID:37127596. PMCID:PMC10150520.
PMID: 37127596
PMCID: PMC10150520
Funding: - H2020 Marie Skłodowska-Curie Actions: 896172
- EMBO: 4431
- rantová Agentura České Republiky: 23-04260L