JUMPER
JUMPER assembles discontinuous transcripts and estimates their abundances from paired-end short-read RNA sequencing alignments to characterize subgenomic RNAs produced by Nidovirales, including SARS-CoV-2, SARS-CoV-1, and MERS-CoV.
Key Features:
- Discontinuous transcript assembly: Solves the discontinuous transcript assembly problem by identifying transcripts T and abundances c given an alignment R of paired-end short reads.
- Statistical model: Employs a maximum likelihood model that accounts for varying transcript lengths.
- Segment graph representation: Represents genomes with a segment graph, a directed acyclic graph (DAG) characterized by a unique Hamiltonian path, distinct from splice graphs used in alternative splicing.
- Biological modeling: Models transcripts arising from discontinuous transcription mediated by the viral RNA-dependent RNA polymerase producing subgenomic RNAs.
- Solution characterization: Provides a compact characterization of solutions as subsets of non-overlapping edges in the segment graph.
- Optimization heuristic: Implements a progressive heuristic that formulates assembly as mixed integer linear programming (MILP) optimization.
- Validation and detection: Recovers canonical and non-canonical transcripts with support from long-read data, recurrence across independent samples, or conserved core sequences.
- Comparative performance: Demonstrates superior performance to existing classical transcript assembly methods in simulations.
- Treatment analysis: Enables analysis of viral transcriptomes under treatment to detect transcript-level viral drug responses.
Scientific Applications:
- Nidovirales transcriptome characterization: Characterizes discontinuous transcriptomes of Nidovirales, including SARS-CoV-2, SARS-CoV-1, and MERS-CoV.
- Subgenomic RNA identification: Identifies subgenomic RNAs and quantifies their abundances produced by discontinuous transcription.
- Non-canonical transcript discovery: Predicts non-canonical transcripts and corroborates them via long-read support, sample recurrence, or conserved core sequences.
- Method benchmarking: Benchmarks assembly accuracy against existing classical transcript assembly methods using simulations.
- Drug response detection: Detects transcript-level changes in viral gene expression under treatment to assess drug responses.
Methodology:
Identifies transcripts T and abundances c from an alignment R of paired-end short reads using a maximum likelihood model that accommodates variable transcript lengths, represents the genome as a segment graph (a DAG with a unique Hamiltonian path), characterizes solutions as subsets of non-overlapping edges, and optimizes assembly via a progressive heuristic formulated as mixed integer linear programming (MILP).
Topics
Details
- License:
- MIT
- Cost:
- Free of charge (with restrictions)
- Tool Type:
- command-line tool
- Operating Systems:
- Mac, Linux, Windows
- Programming Languages:
- Python, R
- Added:
- 11/11/2021
- Last Updated:
- 11/11/2021
Operations
Publications
Sashittal P, Zhang C, Peng J, El-Kebir M. JUMPER Enables Discontinuous Transcript Assembly in Coronaviruses. Unknown Journal. 2021. doi:10.21203/rs.3.rs-600334/v1.