CanPredict

CanPredict predicts whether specific protein sequence changes (missense variants) are likely to be cancer-associated by integrating multiple computational scores and a classifier to distinguish driver mutations from polymorphisms.


Key Features:

  • Sorting Intolerant From Tolerant (SIFT): Assesses whether an amino acid substitution affects protein function using sequence conservation.
  • Pfam-based LogR.E-value metric: Uses Pfam domain information and a LogR.E-value metric to evaluate the functional impact of mutations within protein domains.
  • Gene Ontology Similarity Score (GOSS): Measures similarity between a gene harboring a variant and known cancer-causing genes based on Gene Ontology annotations.
  • Random forest classifier: Integrates scores from SIFT, Pfam LogR.E-value, and GOSS to predict whether a mutation is likely cancer-associated.
  • Low false-positive rate: Demonstrates a low false-positive rate when distinguishing true cancer-associated mutations from polymorphisms and other missense changes.
  • Prediction of activating and inactivating mutations: Predicts both activating and inactivating oncogenic missense mutations as deleterious.
  • Distinguishes drivers from passengers: Identifies features that differentiate oncogenic (driver) mutations from passenger variants and common SNPs.
  • Utility in large-scale analyses: Applied to large-scale genomic datasets to prioritize recurrently observed mutations that are more likely cancer-associated than rare variants.
  • Identification of novel germline changes: Has identified previously unknown germline variants with potential oncogenic significance, including P1104A in TYK2.

Scientific Applications:

  • Mutation prioritization for experimental validation: Ranks missense variants to guide selection of candidates for functional studies.
  • Driver versus passenger classification: Supports discrimination of causal cancer-associated mutations from passenger variants and polymorphisms.
  • Oncogenomics and translational research: Aids identification of candidate cancer-associated mutations to inform studies of tumorigenesis and potential therapeutic targets.

Methodology:

Computes SIFT, a Pfam-based LogR.E-value metric, and a Gene Ontology Similarity Score (GOSS) for each missense variant and combines these scores using a random forest classifier to predict cancer-association.

Topics

Details

Tool Type:
web application
Added:
2/10/2017
Last Updated:
12/10/2018

Operations

Publications

Kaminker JS, et al. CanPredict: a computational tool for predicting cancer-associated missense mutations. Nucleic Acids Res. 2007; 35:W595-8. doi: 10.1093/nar/gkm405

PMID: 17537827

Kaminker JS, et al. Distinguishing cancer-associated missense mutations from common polymorphisms. Cancer Res. 2007; 67:465-73. doi: 10.1158/0008-5472.CAN-06-1736

PMID: 17234753