Complet_plus

Complet_plus enhances the completeness of protein sequence clustering while maintaining homogeneity to produce more biologically representative clusters.


Key Features:

  • Balancing Homogeneity and Completeness: Complet_plus merges closely related clusters based on verified structural relationships defined by the SCOPe classification to improve completeness while preserving homogeneity.
  • Scalability: The method demonstrates linear scalability in runtime relative to the number of clusters and is applicable to large datasets such as a COG dataset containing over 3 million sequences.
  • Integration with Existing Tools: Complet_plus functions as a post-processing step for clustering outputs from MMseqs2's clusterupdate, CD-HIT, and linclust.
  • Quantitative Improvements: Application with MMseqs2's clusterupdate increased V-measure by 0.09 at the SCOPe superfamily level and 0.05 at the family level and produced substantial increases in Adjusted Mutual Information (AMI) and Adjusted Rand Index (ARI) against biological classifications.

Scientific Applications:

  • Proteome-scale clustering: Improves biological representativeness of clusters in large-scale proteomic datasets.
  • Structural classification benchmarking: Enhances mapping between sequence clusters and structural classifications such as SCOPe for validating clustering accuracy.
  • Functional and evolutionary inference: Supports more accurate inference of protein functions and evolutionary relationships from clustered sequences.

Methodology:

Complet_plus operates as a post-processing step that merges clusters based on verified structural relationships (SCOPe) and demonstrates linear runtime scaling relative to the number of clusters.

Topics

Details

License:
Not licensed
Tool Type:
command-line tool
Programming Languages:
Shell, Python
Added:
3/9/2023
Last Updated:
11/24/2024

Operations

Publications

Nguyen R, Sokhansanj BA, Polikar R, Rosen GL. Complet+: a computationally scalable method to improve completeness of large-scale protein sequence clustering. PeerJ. 2023;11:e14779. doi:10.7717/peerj.14779. PMID:36785708. PMCID:PMC9921987.

PMID: 36785708
PMCID: PMC9921987
Funding: - NSF grants: #1919691, #1936782, #1936791, #2107108