pysradb

pysradb provides programmatic access to NCBI Sequence Read Archive (SRA) metadata and raw next-generation sequencing datasets, enabling systematic querying and retrieval for reproducible analyses.


Key Features:

  • Programmatic querying: Enables scripted queries of SRA metadata to locate datasets and associated metadata entries.
  • Metadata-driven retrieval: Utilizes the curated SRAdb metadata database to search and filter SRA records.
  • Raw sequencing data retrieval: Supports retrieval and downloading of raw sequencing datasets from the SRA.
  • Large-scale dataset handling: Facilitates systematic access to large volumes of next-generation sequencing data.

Scientific Applications:

  • Dataset discovery and reanalysis: Locating and retrieving public next-generation sequencing datasets for reanalysis of experiments.
  • Reproducibility and meta-analysis: Aggregating SRA metadata and raw data to reproduce published analyses and perform cross-study comparisons.
  • Large-scale data acquisition for bioinformatics analyses: Collecting bulk sequencing datasets to support hypothesis testing and downstream computational workflows.

Methodology:

Interfaces with the SRAdb curated metadata database to execute queries that return SRA metadata entries and dataset identifiers enabling retrieval of raw sequencing data.

Topics

Details

License:
BSD-3-Clause
Maturity:
Mature
Cost:
Free of charge
Tool Type:
library
Operating Systems:
Linux, Windows, Mac
Programming Languages:
Python
Added:
5/25/2019
Last Updated:
6/16/2020

Operations

Publications

Choudhary S. pysradb: A Python package to query next-generation sequencing metadata and data from NCBI Sequence Read Archive. F1000Research. 2019;8:532. doi:10.12688/f1000research.18676.1. PMID:31114675. PMCID:PMC6505635.