Readscan
Readscan identifies non-host sequences in high-throughput sequencing datasets to detect potential pathogens and estimate their relative genomic abundance.
Key Features:
- Scalability: Readscan operates as a parallel program that distributes computational tasks across multiple compute-cluster nodes to process large sequencing datasets.
- Parallel processing: The implementation analyzes multiple sequence reads simultaneously to increase throughput.
- Speed: Demonstrated on a simulated dataset of 20.1 million reads, Readscan classified human and viral sequences in under 27 minutes on a Beowulf compute cluster with 16 nodes.
- Accuracy: The tool precisely identifies non-host sequences to support reliable detection of potential pathogen-origin reads.
- Relative abundance estimation: Readscan provides estimates of the relative abundance of genomes from potential pathogens within sequencing samples.
Scientific Applications:
- Metagenomic studies: Distinguishing host and non-host sequences to characterize microbial composition in mixed samples.
- Pathogen detection and surveillance: Rapid identification of potential pathogenic sequences from high-throughput sequencing data.
- Microbial community profiling: Estimating relative genome abundances to compare community structure across samples.
- Epidemiology and environmental microbiology: Supporting analyses for infectious disease investigations and environmental sequencing surveys.
Methodology:
Parallel processing of sequencing reads by distributing computational tasks across multiple compute-cluster nodes for simultaneous analysis of reads.
Topics
Details
- Tool Type:
- command-line tool
- Operating Systems:
- Linux, Mac
- Added:
- 8/3/2017
- Last Updated:
- 11/25/2024
Operations
Publications
Naeem R, Rashid M, Pain A. READSCAN: a fast and scalable pathogen discovery program with accurate genome relative abundance estimation. Bioinformatics. 2012;29(3):391-392. doi:10.1093/bioinformatics/bts684. PMID:23193222. PMCID:PMC3562070.