mGSZ

mGSZ applies the Gene Set Z-score (GSZ) scoring function with asymptotic P-value estimation to perform gene set analysis that evaluates the collective effects of gene sets on biological processes.


Key Features:

  • Asymptotic P-Value Calculation: Uses asymptotic P-values instead of extensive permutation-based empirical P-values to reduce the need for numerous permutations and accelerate computation.
  • GSZ Scoring Function: Centralizes the GSZ scoring function, which was demonstrated to outperform seven other popular gene set scoring functions in comparative studies.
  • Max-Mean Statistics Replacement: Replaces max-mean statistics with the GSZ scoring function to enhance gene set analysis performance.
  • Permutation Importance: Emphasizes the necessity of both gene and sample permutations for accurate P-value estimation and highlights risks when either permutation is omitted.
  • Comparison with Empirical Methods: Employs an asymptotic approach that offers advantages over empirical methods in efficiency and precision of P-value estimation.
  • Performance Evaluation: Demonstrates superior performance in two distinct evaluations and reports that rotation tests do not improve the asymptotic P-values from its methodology.
  • Asymptotic Distribution Models: Proposes well-known asymptotic distribution models for three of the compared gene set analysis methods.
  • R Package Implementation: Provides an implementation as an R package.

Scientific Applications:

  • Genomics: Enables gene set analysis for genomic studies to assess coordinated gene effects on biological processes.
  • Transcriptomics: Supports analysis of transcriptomic data to detect pathway-level changes using GSZ scoring and asymptotic P-values.
  • Systems Biology: Facilitates systems-level interpretation of gene set effects within biological networks and processes.
  • High-Throughput Studies: Suited for high-throughput datasets where reduced computational demands for P-value estimation are beneficial.

Methodology:

Applies the GSZ scoring function and computes asymptotic P-values rather than empirical permutation P-values; substitutes max-mean statistics with GSZ; highlights the use of both gene and sample permutations for P-value accuracy; compares GSZ against seven other scoring functions and evaluates performance in two distinct tests, reporting that rotation tests do not improve asymptotic P-values; proposes asymptotic distribution models for three compared methods; implemented as an R package.

Topics

Details

Tool Type:
command-line tool
Operating Systems:
Linux, Windows, Mac
Programming Languages:
R
Added:
8/3/2017
Last Updated:
11/25/2024

Operations

Publications

Mishra P, Törönen P, Leino Y, Holm L. Gene set analysis: limitations in popular existing methods and proposed improvements. Bioinformatics. 2014;30(19):2747-2756. doi:10.1093/bioinformatics/btu374. PMID:24903419.

Documentation

Links