POSSUM
POSSUM generates numerical sequence feature descriptors derived from Position-Specific Scoring Matrices (PSSMs) for use in machine-learning analyses of protein sequences and protein attribute prediction.
Key Features:
- PSSM-based feature generation: Produces numerical sequence feature descriptors derived from Position-Specific Scoring Matrices (PSSMs) that encode evolutionary information in protein sequences.
- Twenty-one descriptor types: Generates 21 distinct types of PSSM-based feature descriptors.
- Machine-learning compatibility: Outputs numerical descriptors suitable for feature extraction, feature selection, and benchmarking in machine-learning models.
- Protein attribute relevance: Provides descriptors specifically intended to support predictors of protein attributes.
Scientific Applications:
- Feature extraction for ML: Supplying PSSM-derived numerical descriptors to train classifiers and regressors on protein sequence data.
- Feature selection and benchmarking: Enabling selection and benchmarking of PSSM-based features during model development and evaluation.
- Protein attribute prediction: Enhancing the input feature set for predictors of protein attributes by incorporating evolutionary information from PSSMs.
Methodology:
Generates numerical sequence feature descriptors from Position-Specific Scoring Matrices (PSSMs), producing 21 distinct PSSM-based descriptor types.
Topics
Details
- Tool Type:
- web application
- Operating Systems:
- Linux, Windows, Mac
- Added:
- 6/7/2018
- Last Updated:
- 11/25/2024
Operations
Publications
Wang J, Yang B, Revote J, Leier A, Marquez-Lago TT, Webb G, Song J, Chou K, Lithgow T. POSSUM: a bioinformatics toolkit for generating numerical sequence feature descriptors based on PSSM profiles. Bioinformatics. 2017;33(17):2756-2758. doi:10.1093/bioinformatics/btx302. PMID:28903538.