PROSE
PROSE computes interpretable enrichment scores from observed proteins using gene co-expression matrices to recover missing proteins and associate proteomic profiles with phenotypes.
Key Features:
- Scalability and Efficiency: Scalable, computationally efficient pipeline for processing large proteomic datasets.
- Enrichment Scoring: Computes interpretable and stable enrichment scores from an observed protein set that quantify each protein's relative importance to specific phenotypes.
- Missing Protein Recovery: Leverages gene co-expression matrices to reproducibly recover missing proteins from proteomic data, improving proteome coverage.
- Phenotype Association: Aligns enrichment scores with source phenotypes to relate protein presence or absence to phenotypic characteristics.
Scientific Applications:
- Proteomic dataset reanalysis (e.g., CCLE): Applied to reanalyze proteomic datasets such as the Cancer Cell Line Encyclopedia (CCLE) to enhance protein coverage and interpretation.
- Oncogenic dependency prediction: Used to predict oncogenic dependencies from proteomic profiles.
- Regulatory module identification: Used to identify well-defined regulatory modules from proteomic data.
Methodology:
Uses an input set of observed proteins and gene co-expression matrices to compute enrichment scores that quantify protein importance relative to phenotypes and enable recovery of missing proteins.
Topics
Details
- License:
- MIT
- Cost:
- Free of charge
- Tool Type:
- library
- Operating Systems:
- Mac, Linux, Windows
- Programming Languages:
- Python
- Added:
- 3/31/2022
- Last Updated:
- 3/31/2022
Operations
Publications
Wong BJH, Kong W, Goh WWB. Single-sample proteome enrichment enables missing protein recovery and phenotype association. Unknown Journal. 2021. doi:10.1101/2021.11.13.468488.