BioSamples
BioSamples stores and manages metadata about biological samples used in sequencing, gene expression, and proteomics experiments to enable consistent recording and cross-database linking of sample information.
Key Features:
- Consistent Sample Information Recording: Allows submitters to describe a sample once and reference it in subsequent submissions to assay databases to reduce redundancy and ensure metadata consistency.
- Cross-Database Queries: Supports attribute-based queries (sample types, disease names, sample providers) to enable searches across linked datasets.
- Reference Samples: Maintains a collection of reference samples, including cell lines, and assigns accession numbers for cross-database referencing and interoperability with NCBI and DDBJ.
- Integration with ENA, ArrayExpress, and PRIDE: Links sample metadata to ENA, ArrayExpress, and PRIDE to support consistent sample annotation across these resources.
- Technical Enhancements and Scalability: Implements improvements that enable scaling to broader communities and more diverse submissions while enriching metadata content and supporting FAIR principles.
- Technical Infrastructure: Provides infrastructure to support engagement with initiatives such as FAIRplus and the Global Alliance for Genomics and Health.
Scientific Applications:
- Standardization of experimental metadata: Standardizes sample metadata for sequencing, gene expression, and proteomics experiments to enable consistent downstream analysis.
- Cross-database sample discovery: Enables discovery across ENA, ArrayExpress, and PRIDE by sample characteristics such as disease or provider.
- Reference sample reuse: Provides accessioned reference cell lines to facilitate reuse and comparability across studies.
- Genotype–phenotype linking in plants: Facilitates linking genotypic to phenotypic information within plant research communities.
- Multi-omics integration for disease research: Supports complex multi-omics data integration relevant to COVID-19 research.
- Data management use cases: Supports exemplar initiatives such as the ReSOLUTE project to increase sample findability and improve data management practices.
Methodology:
Assigns accession numbers, records single-instance sample descriptions with referencing for subsequent submissions, supports attribute-based querying (sample types, disease names, sample providers), links metadata across ENA, ArrayExpress, PRIDE, NCBI and DDBJ, and applies metadata enrichment and scalability improvements to support FAIR principles.
Topics
Details
- Tool Type:
- web application
- Operating Systems:
- Linux, Windows, Mac
- Added:
- 2/4/2015
- Last Updated:
- 3/28/2022
Operations
Publications
Courtot M, Gupta D, Liyanage I, Xu F, Burdett T. BioSamples database: FAIRer samples metadata to accelerate research data management. Nucleic Acids Research. 2021;50(D1):D1500-D1507. doi:10.1093/nar/gkab1046. PMID:34747489. PMCID:PMC8728232.
Gostev M, Faulconbridge A, Brandizi M, Fernandez-Banet J, Sarkans U, Brazma A, Parkinson H. The BioSample Database (BioSD) at the European Bioinformatics Institute. Nucleic Acids Research. 2011;40(D1):D64-D70. doi:10.1093/nar/gkr937. PMID:22096232. PMCID:PMC3245134.