Seq2Fun
Seq2Fun performs ultra-fast functional profiling of RNA-seq reads to quantify gene and pathway abundance for experiments including non-model organisms without reference genomes.
Key Features:
- Ultra-Fast Processing: Processes over 2 million reads per minute and is reported to be more than 120 times faster than workflows relying on de novo transcriptome assembly.
- Direct Functional Quantification: Performs functional quantification directly on reads while bypassing de novo transcriptome assembly and reducing file I/O.
- Comprehensive Pipeline: Implements raw read quality control including sequencing error correction, poly(A) tail removal, and joining overlapped paired-end reads; translates reads into all possible amino acid fragments and performs DNA-to-protein searches against a curated protein database; and produces gene abundance tables, pathway and species hit tables, an HTML report, and annotated clean reads.
- Efficiency on Standard Hardware: Implemented in C++ and optimized to run on personal computers with limited CPU and memory resources.
Scientific Applications:
- Functional profiling of non-model organisms: Enables rapid gene expression analysis, pathway elucidation, and species identification from RNA-seq data when reference genomes are unavailable and generates outputs for downstream analyses.
Methodology:
Performs raw read quality control (sequencing error correction, poly(A) removal, joining overlapped paired-end reads), translates each read into all possible amino acid fragments, conducts DNA-to-protein homology searches against a curated protein database, and directly quantifies functional features from read-level matches; implemented in C++ and optimized for limited CPU and memory.
Topics
Details
- Tool Type:
- command-line tool, web application
- Programming Languages:
- C, C++
- Added:
- 11/29/2021
- Last Updated:
- 11/29/2021
Operations
Publications
Liu P, Ewald J, Galvez JH, Head J, Crump D, Bourque G, Basu N, Xia J. Ultrafast functional profiling of RNA-seq data for nonmodel organisms. Genome Research. 2021;31(4):713-720. doi:10.1101/gr.269894.120. PMID:33731361. PMCID:PMC8015844.