contamDE-lm
contamDE-lm applies a novel linear model to next-generation RNA-seq data to perform differential gene expression analysis while adjusting for cellular contamination in paired tumor–normal samples.
Key Features:
- Accounting for Cellular Contamination: Accounts for cellular contamination in tumor RNA-seq by incorporating gene-wise information to reduce individual residual variances and mitigate contamination-driven false positives and false negatives.
- Novel Linear Model: Implements a linear-model framework that addresses tumor–normal correlations in paired samples to mitigate bias and improve reliability of DE results.
- Computational Efficiency: Provides computational speed advantages for large RNA-seq datasets with comparisons noted against limma and DESeq2.
- Statistical Robustness: Manages contamination-induced noise to improve detection of differentially expressed genes between normal and tumor samples and among tumor subtypes.
- Sequencing and Design Suitability: Designed for next-generation RNA-seq data and paired tumor–normal study designs.
Scientific Applications:
- Cancer differential expression analysis: Analyzes differential gene expression between tumor and adjacent normal tissues to provide insights into tumorigenesis and potential therapeutic targets.
- Heterogeneous tissue studies: Applies to RNA-seq datasets from heterogeneous or contaminated tissue samples to improve DE inference in presence of mixed cell populations.
Methodology:
Applies a novel linear-model approach that integrates gene-wise information to adjust for cellular contamination, reduce residual variances, and account for tumor–normal correlations in paired RNA-seq differential expression analyses; it has been applied to simulated data and to TCGA and GEO datasets.
Topics
Details
- Tool Type:
- library
- Programming Languages:
- R
- Added:
- 1/18/2021
- Last Updated:
- 3/11/2021
Operations
Publications
Ji Y, Yu C, Zhang H. contamDE-lm: linear model-based differential gene expression analysis using next-generation RNA-seq data from contaminated tumor samples. Bioinformatics. 2020;36(8):2492-2499. doi:10.1093/bioinformatics/btaa006. PMID:31917401.