UniqueProt
UniqueProt generates representative sets of protein sequences by removing redundancy based on HSSP-value sequence similarity to produce maximally non-redundant datasets for downstream protein sequence analyses.
Key Features:
- Redundancy Elimination: Systematically removes redundant protein sequences to produce a non-redundant representative set.
- HSSP-value Utilization: Employs HSSP-values to assess pairwise sequence similarity for redundancy decisions.
- Greedy Algorithm Approach: Uses a straightforward greedy selection algorithm to identify representative sequences.
- Representative Set Maximization: Aims to produce the largest possible representative datasets without inherent bias.
Scientific Applications:
- Data Set Bias Reduction: Generates unbiased protein sequence datasets to improve integrity and representativeness in analyses.
- Protein Sequence Analysis: Produces non-redundant sequence sets for phylogenetic analysis, functional annotation, and evolutionary studies.
Methodology:
Applies a greedy selection algorithm that evaluates sequence similarity using HSSP-values and selects sequences that best represent dataset diversity rather than placing representatives at cluster centers, with selection guided by problem-specific criteria.
Topics
Details
- Tool Type:
- web application
- Operating Systems:
- Linux
- Programming Languages:
- Perl
- Added:
- 2/10/2017
- Last Updated:
- 11/25/2024
Operations
Publications
Mika S. UniqueProt: creating representative protein sequence sets. Nucleic Acids Research. 2003;31(13):3789-3791. doi:10.1093/nar/gkg620. PMID:12824419. PMCID:PMC169026.