POSSUM

POSSUM generates numerical sequence feature descriptors derived from Position-Specific Scoring Matrices (PSSMs) for use in machine-learning analyses of protein sequences and protein attribute prediction.


Key Features:

  • PSSM-based feature generation: Produces numerical sequence feature descriptors derived from Position-Specific Scoring Matrices (PSSMs) that encode evolutionary information in protein sequences.
  • Twenty-one descriptor types: Generates 21 distinct types of PSSM-based feature descriptors.
  • Machine-learning compatibility: Outputs numerical descriptors suitable for feature extraction, feature selection, and benchmarking in machine-learning models.
  • Protein attribute relevance: Provides descriptors specifically intended to support predictors of protein attributes.

Scientific Applications:

  • Feature extraction for ML: Supplying PSSM-derived numerical descriptors to train classifiers and regressors on protein sequence data.
  • Feature selection and benchmarking: Enabling selection and benchmarking of PSSM-based features during model development and evaluation.
  • Protein attribute prediction: Enhancing the input feature set for predictors of protein attributes by incorporating evolutionary information from PSSMs.

Methodology:

Generates numerical sequence feature descriptors from Position-Specific Scoring Matrices (PSSMs), producing 21 distinct PSSM-based descriptor types.

Topics

Details

Tool Type:
web application
Operating Systems:
Linux, Windows, Mac
Added:
6/7/2018
Last Updated:
11/25/2024

Operations

Publications

Wang J, Yang B, Revote J, Leier A, Marquez-Lago TT, Webb G, Song J, Chou K, Lithgow T. POSSUM: a bioinformatics toolkit for generating numerical sequence feature descriptors based on PSSM profiles. Bioinformatics. 2017;33(17):2756-2758. doi:10.1093/bioinformatics/btx302. PMID:28903538.

PMID: 28903538
Funding: - NHMRC: 1092262

Documentation