Genomic benchmarks

Genomic benchmarks provides a curated collection of genomic sequence datasets and baseline deep learning resources for benchmarking classification of regulatory elements such as promoters, enhancers, and open chromatin regions across human, mouse, and roundworm.


Key Features:

  • Curated Datasets: Nine curated datasets focused on regulatory elements (promoters, enhancers, and open chromatin regions) from human, mouse, and roundworm.
  • Baseline Model: A simple convolutional neural network (CNN) is provided as a baseline model for performance comparison.
  • Integration with Deep Learning Libraries: An interface compatible with commonly used deep learning libraries is provided to enable integration into existing workflows.
  • Training Framework: A training framework is implemented to support reproducible and comparable machine learning experiments.

Scientific Applications:

  • Genome Annotation: Use curated datasets to improve annotation of genomic regions with regulatory function.
  • Identification of Functional Elements: Benchmark classification methods for detecting promoters, enhancers, and open chromatin regions.
  • Machine Learning Research in Genomics: Provide standardized datasets and a baseline model to enable comparative evaluation of deep learning methods.

Methodology:

Datasets were constructed by mining publicly available databases and integrating existing data from published articles.

Topics

Details

License:
Apache-2.0
Cost:
Free of charge
Tool Type:
library
Operating Systems:
Mac, Linux, Windows
Programming Languages:
Python
Added:
1/10/2024
Last Updated:
11/24/2024

Operations

Publications

Grešová K, Martinek V, Čechák D, Šimeček P, Alexiou P. Genomic benchmarks: a collection of datasets for genomic sequence classification. BMC Genomic Data. 2023;24(1). doi:10.1186/s12863-023-01123-8. PMID:37127596. PMCID:PMC10150520.

PMID: 37127596
Funding: - H2020 Marie Skłodowska-Curie Actions: 896172 - EMBO: 4431 - rantová Agentura České Republiky: 23-04260L