GINCCo
GINCCo constructs unsupervised computational graph models constrained by gene interaction networks (e.g., protein-protein interaction (PPI) networks) to integrate biological knowledge into predictive modeling and interpretation of gene expression data for applications such as cancer phenotype prediction.
Key Features:
- Unsupervised Computational Graph Construction: Automates creation of computational graph models from gene expression data without manual feature engineering.
- Structural Inductive Biases from Biological Knowledge: Embeds structures from gene interaction graphs such as protein-protein interaction (PPI) networks to guide model construction.
- Topological Clustering Algorithms: Applies topological clustering on PPI networks to derive network-derived structures for model nodes.
- Biologically Relevant Entities Representation: Represents genes, candidate protein complexes, and phenotypes as distinct entities within the computational graphs.
- Model Regularization and Performance: Structures models around biologically meaningful interactions to achieve strong predictive performance with substantially fewer parameters than support vector machines (SVMs) or fully connected multi-layer perceptrons (MLPs).
- Post-hoc Enrichment Analyses: Facilitates guided post-hoc enrichment analyses to identify influential gene sets associated with specific phenotypes.
Scientific Applications:
- Cancer phenotype prediction: Uses PPI-constrained computational graphs to predict cancer-related phenotypes from gene expression data, demonstrating improved performance relative to conventional models.
- Gene set interpretation via enrichment analysis: Supports identification and interpretation of phenotype-associated gene sets through guided post-hoc enrichment analyses.
- Protein complex discovery and representation: Leverages topological clustering on PPI networks to identify candidate protein complexes that serve as model entities.
Methodology:
Unsupervised construction of computational graph models constrained by prior knowledge from gene interaction networks using topological clustering algorithms on PPI networks, with graphs that explicitly represent genes, candidate protein complexes, and phenotypes and support guided post-hoc enrichment analyses.
Topics
Details
- License:
- MIT
- Cost:
- Free of charge
- Tool Type:
- library
- Operating Systems:
- Mac, Linux, Windows
- Programming Languages:
- Python
- Added:
- 5/27/2022
- Last Updated:
- 5/27/2022
Operations
Publications
Scherer P, Trębacz M, Simidjievski N, Viñas R, Shams Z, Terre HA, Jamnik M, Liò P. Unsupervised construction of computational graphs for gene expression data with explicit structural inductive biases. Bioinformatics. 2021;38(5):1320-1327. doi:10.1093/bioinformatics/btab830. PMID:34888618. PMCID:PMC8826027.