SpiderSeqR

SpiderSeqR enables unified discovery and retrieval of multi-omic datasets and metadata from NCBI GEO and SRA to support reuse and analysis of genomic, transcriptomic, translatomic, and epigenomic data generated by massively parallel sequencing technologies.


Key Features:

  • Integration with Major Databases: Integrates programmatic access to NCBI GEO and SRA to expose datasets and their associated annotations.
  • Unified Search and Accession Conversion: Performs integrated searches across SRA and GEO and converts between different database accession identifiers to map records between repositories.
  • Advanced Filtering and Query Management: Applies user-specified filtering criteria to search results and supports saving and reusing past queries for repeated retrieval tasks.
  • Support for Multi-omic Sequencing Data: Targets genomic, transcriptomic, translatomic, and epigenomic datasets produced by massively parallel sequencing technologies to facilitate dataset discovery and linkage of annotations.

Scientific Applications:

  • Genomic and Multi-omic Research: Enables retrieval and aggregation of genomic, transcriptomic, translatomic, and epigenomic datasets for downstream bioinformatic analyses.
  • Support for Large-scale Projects: Assists data discovery and reuse for large consortia and initiatives such as the Human Genome Project and ENCODE by linking repository records and annotations.
  • Data Repurposing and Meta-analysis: Facilitates repurposing of public SRA and GEO datasets for novel research questions and cross-study meta-analyses by improving metadata retrieval and mapping.

Methodology:

Implements a web-crawling framework that automates searching and filtering of metadata from SRA and GEO and performs accession identifier conversion to map records between repositories.

Topics

Details

Programming Languages:
R
Added:
1/18/2021
Last Updated:
2/21/2021

Operations

Publications

Sozanska AM, Fletcher C, Bihary D, Samarajiwa SA. SpiderSeqR: an R package for crawling the web of high-throughput multi-omic data repositories for data-sets and annotation. Unknown Journal. 2020. doi:10.1101/2020.04.13.039420.