PROSE

PROSE computes interpretable enrichment scores from observed proteins using gene co-expression matrices to recover missing proteins and associate proteomic profiles with phenotypes.


Key Features:

  • Scalability and Efficiency: Scalable, computationally efficient pipeline for processing large proteomic datasets.
  • Enrichment Scoring: Computes interpretable and stable enrichment scores from an observed protein set that quantify each protein's relative importance to specific phenotypes.
  • Missing Protein Recovery: Leverages gene co-expression matrices to reproducibly recover missing proteins from proteomic data, improving proteome coverage.
  • Phenotype Association: Aligns enrichment scores with source phenotypes to relate protein presence or absence to phenotypic characteristics.

Scientific Applications:

  • Proteomic dataset reanalysis (e.g., CCLE): Applied to reanalyze proteomic datasets such as the Cancer Cell Line Encyclopedia (CCLE) to enhance protein coverage and interpretation.
  • Oncogenic dependency prediction: Used to predict oncogenic dependencies from proteomic profiles.
  • Regulatory module identification: Used to identify well-defined regulatory modules from proteomic data.

Methodology:

Uses an input set of observed proteins and gene co-expression matrices to compute enrichment scores that quantify protein importance relative to phenotypes and enable recovery of missing proteins.

Topics

Details

License:
MIT
Cost:
Free of charge
Tool Type:
library
Operating Systems:
Mac, Linux, Windows
Programming Languages:
Python
Added:
3/31/2022
Last Updated:
3/31/2022

Operations

Publications

Wong BJH, Kong W, Goh WWB. Single-sample proteome enrichment enables missing protein recovery and phenotype association. Unknown Journal. 2021. doi:10.1101/2021.11.13.468488.