SpliceDB

SpliceDB integrates RNA sequencing (RNA-seq) data with mass spectrometry-based proteogenomics to construct compact splice graph databases from RNA-seq reads for improved genomic annotation.


Key Features:

  • Data compression and efficiency: Compresses aligned RNA-seq SAM files into a FASTA-formatted splice graph database achieving up to 1000× reduction in size (e.g., C. elegans: 496.2 GB to 410 MB) while maintaining sensitivity.
  • Retention of read information: Constructs a compact database that retains all essential information expressed in RNA-seq reads.
  • Automated pipeline: Implements a fully automated pipeline for integrating transcriptomic and mass spectrometry-based proteogenomic data.
  • Proteogenomic validation: Integrates mass spectrometry evidence to validate splicing events and translated regions for genomic annotation.
  • Novel event identification: Identified 4,044 novel events in a custom dataset, including 215 novel genes, 808 novel exons, 12 alternative splicing events, 618 gene-boundary corrections, 245 exon-boundary changes, 938 frame shifts, 1,166 reverse strands, and 42 translated untranslated regions (UTRs).

Scientific Applications:

  • Genomic annotation improvement: Improves genome annotations by combining RNA-seq-derived splice graphs with proteomic evidence.
  • Splicing and translation confirmation: Confirms splicing events and transcribed regions at the protein level using mass spectrometry-based proteogenomics.
  • Gene expression and splicing analysis: Enables integrated analysis of gene expression and alternative splicing across diverse cellular conditions.

Methodology:

Constructs a compact splice graph database from RNA-seq reads, compresses aligned SAM files into a FASTA-formatted splice graph database, and integrates mass spectrometry-based proteogenomics via an automated pipeline.

Topics

Collections

Details

Tool Type:
command-line tool
Operating Systems:
Linux, Windows, Mac
Programming Languages:
Java
Added:
12/18/2017
Last Updated:
11/25/2024

Operations

Publications

Woo S, Cha SW, Merrihew G, He Y, Castellana N, Guest C, MacCoss M, Bafna V. Proteogenomic Database Construction Driven from Large Scale RNA-seq Data. Journal of Proteome Research. 2013;13(1):21-28. doi:10.1021/pr400294c. PMID:23802565. PMCID:PMC4034692.

Documentation

Links