UProC

UProC classifies protein sequences using the Mosaic Matching algorithm to enable rapid functional classification of large-scale metagenomic and proteomic sequence datasets.


Key Features:

  • Novel Algorithm: UProC implements the Mosaic Matching algorithm for rapid and efficient protein sequence classification.
  • Performance Efficiency: The method is reported to be up to three orders of magnitude faster than traditional profile-based methods.
  • Increased Sensitivity: In simulated metagenome studies, UProC demonstrated up to 80% higher sensitivity on unassembled short reads (100 bp) compared to existing methods.

Scientific Applications:

  • Metagenome analysis: Rapid classification of protein sequences from metagenomes and unassembled short reads to profile community functions.
  • Functional genomics: Assignment of proteins to known families to support functional annotation of sequence data.
  • Proteomics: Classification of protein sequences for interpretation of proteomic datasets.
  • Systems biology: Large-scale functional profiling of microbial communities and their protein repertoires.

Methodology:

UProC applies the Mosaic Matching algorithm to compare input sequences against databases of known protein families such as Pfam.

Topics

Details

Tool Type:
command-line tool
Operating Systems:
Linux, Mac
Programming Languages:
C
Added:
8/3/2017
Last Updated:
11/25/2024

Operations

Publications

Meinicke P. UProC: tools for ultra-fast protein domain classification. Bioinformatics. 2014;31(9):1382-1388. doi:10.1093/bioinformatics/btu843. PMID:25540185. PMCID:PMC4410661.

Documentation

Links