UniqueProt

UniqueProt generates representative sets of protein sequences by removing redundancy based on HSSP-value sequence similarity to produce maximally non-redundant datasets for downstream protein sequence analyses.


Key Features:

  • Redundancy Elimination: Systematically removes redundant protein sequences to produce a non-redundant representative set.
  • HSSP-value Utilization: Employs HSSP-values to assess pairwise sequence similarity for redundancy decisions.
  • Greedy Algorithm Approach: Uses a straightforward greedy selection algorithm to identify representative sequences.
  • Representative Set Maximization: Aims to produce the largest possible representative datasets without inherent bias.

Scientific Applications:

  • Data Set Bias Reduction: Generates unbiased protein sequence datasets to improve integrity and representativeness in analyses.
  • Protein Sequence Analysis: Produces non-redundant sequence sets for phylogenetic analysis, functional annotation, and evolutionary studies.

Methodology:

Applies a greedy selection algorithm that evaluates sequence similarity using HSSP-values and selects sequences that best represent dataset diversity rather than placing representatives at cluster centers, with selection guided by problem-specific criteria.

Topics

Details

Tool Type:
web application
Operating Systems:
Linux
Programming Languages:
Perl
Added:
2/10/2017
Last Updated:
11/25/2024

Operations

Publications

Mika S. UniqueProt: creating representative protein sequence sets. Nucleic Acids Research. 2003;31(13):3789-3791. doi:10.1093/nar/gkg620. PMID:12824419. PMCID:PMC169026.

Documentation