RepeatedHighDim

RepeatedHighDim: Statistical framework for high-dimensional gene expression and proteomics analysis

RepeatedHighDim implements global testing procedures to evaluate associations between high-dimensional molecular measurements, including gene expression and proteomics data, and experimental factors.


Key Features:

  • Global Test Procedures: Adapts global tests originally developed for RNA transcript analysis to high-throughput proteomics datasets.
  • Mixed Linear Model: Applies mixed linear models to analyze protein-level data while accounting for incomplete observations.
  • Permutation Procedure: Uses non-parametric permutation testing to assess statistical significance and power in complex proteomics data structures.
  • Imputation Techniques: Implements imputation strategies for missing values, validated by simulation to reduce bias under defined conditions.
  • Gene Ontology (GO) Term Ranking: Ranks GO terms associated with protein sets to evaluate biological processes and pathways.
  • Multiple Protein Representation Handling: Aggregates multiple 2-D gel spots corresponding to a single protein into unified protein sets.

Scientific Applications:

  • Proteomics Association Analysis: Identifies associations between proteins, including calreticulin, and experimental conditions or biological processes in datasets such as heart muscle tissue studies.

Methodology:

Combines mixed linear modeling, permutation-based significance testing, and missing value imputation to maintain robustness and statistical power in high-dimensional proteomics and gene expression datasets.

Topics

Details

Tool Type:
library
Operating Systems:
Linux, Windows, Mac
Programming Languages:
R
Added:
8/3/2017
Last Updated:
11/25/2024

Operations

Publications

Jung K, Dihazi H, Bibi A, Dihazi GH, Beißbarth T. Adaption of the global test idea to proteomics data with missing values. Bioinformatics. 2014;30(10):1424-1430. doi:10.1093/bioinformatics/btu062. PMID:24489372.

Documentation

Links