SpliceDB
SpliceDB integrates RNA sequencing (RNA-seq) data with mass spectrometry-based proteogenomics to construct compact splice graph databases from RNA-seq reads for improved genomic annotation.
Key Features:
- Data compression and efficiency: Compresses aligned RNA-seq SAM files into a FASTA-formatted splice graph database achieving up to 1000× reduction in size (e.g., C. elegans: 496.2 GB to 410 MB) while maintaining sensitivity.
- Retention of read information: Constructs a compact database that retains all essential information expressed in RNA-seq reads.
- Automated pipeline: Implements a fully automated pipeline for integrating transcriptomic and mass spectrometry-based proteogenomic data.
- Proteogenomic validation: Integrates mass spectrometry evidence to validate splicing events and translated regions for genomic annotation.
- Novel event identification: Identified 4,044 novel events in a custom dataset, including 215 novel genes, 808 novel exons, 12 alternative splicing events, 618 gene-boundary corrections, 245 exon-boundary changes, 938 frame shifts, 1,166 reverse strands, and 42 translated untranslated regions (UTRs).
Scientific Applications:
- Genomic annotation improvement: Improves genome annotations by combining RNA-seq-derived splice graphs with proteomic evidence.
- Splicing and translation confirmation: Confirms splicing events and transcribed regions at the protein level using mass spectrometry-based proteogenomics.
- Gene expression and splicing analysis: Enables integrated analysis of gene expression and alternative splicing across diverse cellular conditions.
Methodology:
Constructs a compact splice graph database from RNA-seq reads, compresses aligned SAM files into a FASTA-formatted splice graph database, and integrates mass spectrometry-based proteogenomics via an automated pipeline.
Topics
Collections
Details
- Tool Type:
- command-line tool
- Operating Systems:
- Linux, Windows, Mac
- Programming Languages:
- Java
- Added:
- 12/18/2017
- Last Updated:
- 11/25/2024
Operations
Publications
Woo S, Cha SW, Merrihew G, He Y, Castellana N, Guest C, MacCoss M, Bafna V. Proteogenomic Database Construction Driven from Large Scale RNA-seq Data. Journal of Proteome Research. 2013;13(1):21-28. doi:10.1021/pr400294c. PMID:23802565. PMCID:PMC4034692.