subgraphquant
subgraphquant computes confidence ranges for transcript abundances to assess and quantify non-identifiability in RNA-seq transcript expression quantification.
Key Features:
- Graph Quantification: Generalizes traditional transcript quantification by accounting for incomplete reference transcriptomes and allowing unannotated transcripts to explain observed RNA-seq reads.
- Confidence Range Calculation: Computes a "confidence range of expression" for each transcript representing possible abundance levels across equally optimal estimates.
- Error Analysis: Application to the Human Body Map data indicates that 35%–50% of transcripts may have inaccurate quantification caused by non-identifiability.
- Isoform Comparison: Quantifies that 20%–47% inaccuracy can prevent reliable ranking of expression levels between a transcript and its sibling isoforms within a sample.
- Differential Expression Analysis: Identifies which detected differentially expressed transcripts across RNA-seq samples remain reliable and highlights exceptions affected by non-identifiability.
Scientific Applications:
- Reliability Assessment: Assess the reliability of transcript quantification in the presence of non-identifiable estimates.
- Isoform Expression Comparison: Compare and evaluate the ranking of isoform expression levels within samples or across sample groups while accounting for non-identifiability.
- Differential Expression Validation: Validate differential expression results across RNA-seq samples against potential errors introduced by non-identifiability.
Methodology:
Generalizes transcript quantification to a graph-based formulation that permits unannotated transcripts and computes per-transcript confidence ranges representing possible abundances across equally optimal estimates; applied to datasets such as the Human Body Map to quantify the fraction of affected transcripts.
Topics
Details
- Programming Languages:
- R, C++, Python
- Added:
- 1/14/2020
- Last Updated:
- 1/16/2021
Operations
Publications
Zheng H, Ma C, Kingsford C. Deriving Ranges of Optimal Estimated Transcript Expression Due to Non-identifiability. Unknown Journal. 2019. doi:10.1101/2019.12.13.875625.