PCfun

PCfun predicts Gene Ontology (GO) annotations for protein complexes by combining PubMed Central-derived full-text word embeddings with unsupervised nearest-neighbor retrieval and supervised Random Forest models trained on the CORUM database.


Key Features:

  • Hybrid Machine Learning: Combines unsupervised nearest-neighbor retrieval on PubMed Central full-text word embeddings with supervised Random Forest models.
  • Unsupervised Nearest-Neighbor GO Retrieval: Identifies nearest neighbor Gene Ontology (GO) term vectors for protein complex queries using semantic relationships preserved in the embedding space.
  • Supervised Random Forest Models: Trains Random Forest models on the CORUM protein complex database to recover GO terms associated with protein complexes.
  • Integration via Enrichment Testing: Consolidates predictions by performing statistical tests for the enrichment of top nearest-neighbor GO terms within the child terms predicted by the Random Forest models.
  • PubMed Central-derived Embeddings: Uses word embeddings generated from machine reading of over one million open-access PubMed Central full-text articles.

Scientific Applications:

  • Protein Complex Functional Annotation: Provides GO annotations to interpret the biological processes, cellular components, and molecular functions of protein complexes.
  • Interpretation of Cellular States and Phenotypes: Enables linking protein complex composition to cellular states and phenotypic outcomes via GO terms.
  • Systems Biology and Large-scale Characterization: Facilitates large-scale functional characterization of protein complexes to bridge multi-layer molecular profiling and functional understanding.

Methodology:

Generates word embeddings from machine reading of >1,000,000 open-access PubMed Central full-text articles, applies unsupervised nearest-neighbor retrieval of GO term vectors for protein complex queries, trains Random Forest models on the CORUM database to predict GO terms, and integrates results via statistical enrichment testing of nearest-neighbor terms within child terms predicted by the Random Forest models.

Topics

Details

License:
MIT
Tool Type:
command-line tool
Programming Languages:
Python
Added:
3/19/2021
Last Updated:
3/27/2021

Operations

Publications

Sharma VS, Fossati A, Ciuffa R, Buljan M, Williams EG, Chen Z, Shao W, Pedrioli PG, Purcell AW, Rodríguez Martínez M, Song J, Manica M, Aebersold R, Li C. Towards a systematic characterization of protein complex function: a natural language processing and machine-learning framework. Unknown Journal. 2021. doi:10.1101/2021.02.24.432789.