pysradb
pysradb provides programmatic access to NCBI Sequence Read Archive (SRA) metadata and raw next-generation sequencing datasets, enabling systematic querying and retrieval for reproducible analyses.
Key Features:
- Programmatic querying: Enables scripted queries of SRA metadata to locate datasets and associated metadata entries.
- Metadata-driven retrieval: Utilizes the curated SRAdb metadata database to search and filter SRA records.
- Raw sequencing data retrieval: Supports retrieval and downloading of raw sequencing datasets from the SRA.
- Large-scale dataset handling: Facilitates systematic access to large volumes of next-generation sequencing data.
Scientific Applications:
- Dataset discovery and reanalysis: Locating and retrieving public next-generation sequencing datasets for reanalysis of experiments.
- Reproducibility and meta-analysis: Aggregating SRA metadata and raw data to reproduce published analyses and perform cross-study comparisons.
- Large-scale data acquisition for bioinformatics analyses: Collecting bulk sequencing datasets to support hypothesis testing and downstream computational workflows.
Methodology:
Interfaces with the SRAdb curated metadata database to execute queries that return SRA metadata entries and dataset identifiers enabling retrieval of raw sequencing data.
Topics
Details
- License:
- BSD-3-Clause
- Maturity:
- Mature
- Cost:
- Free of charge
- Tool Type:
- library
- Operating Systems:
- Linux, Windows, Mac
- Programming Languages:
- Python
- Added:
- 5/25/2019
- Last Updated:
- 6/16/2020
Operations
Publications
Choudhary S. pysradb: A Python package to query next-generation sequencing metadata and data from NCBI Sequence Read Archive. F1000Research. 2019;8:532. doi:10.12688/f1000research.18676.1. PMID:31114675. PMCID:PMC6505635.