RepeatedHighDim
RepeatedHighDim: Statistical framework for high-dimensional gene expression and proteomics analysis
RepeatedHighDim implements global testing procedures to evaluate associations between high-dimensional molecular measurements, including gene expression and proteomics data, and experimental factors.
Key Features:
- Global Test Procedures: Adapts global tests originally developed for RNA transcript analysis to high-throughput proteomics datasets.
- Mixed Linear Model: Applies mixed linear models to analyze protein-level data while accounting for incomplete observations.
- Permutation Procedure: Uses non-parametric permutation testing to assess statistical significance and power in complex proteomics data structures.
- Imputation Techniques: Implements imputation strategies for missing values, validated by simulation to reduce bias under defined conditions.
- Gene Ontology (GO) Term Ranking: Ranks GO terms associated with protein sets to evaluate biological processes and pathways.
- Multiple Protein Representation Handling: Aggregates multiple 2-D gel spots corresponding to a single protein into unified protein sets.
Scientific Applications:
- Proteomics Association Analysis: Identifies associations between proteins, including calreticulin, and experimental conditions or biological processes in datasets such as heart muscle tissue studies.
Methodology:
Combines mixed linear modeling, permutation-based significance testing, and missing value imputation to maintain robustness and statistical power in high-dimensional proteomics and gene expression datasets.
Topics
Details
- Tool Type:
- library
- Operating Systems:
- Linux, Windows, Mac
- Programming Languages:
- R
- Added:
- 8/3/2017
- Last Updated:
- 11/25/2024
Operations
Publications
Jung K, Dihazi H, Bibi A, Dihazi GH, Beißbarth T. Adaption of the global test idea to proteomics data with missing values. Bioinformatics. 2014;30(10):1424-1430. doi:10.1093/bioinformatics/btu062. PMID:24489372.