Complet_plus
Complet_plus enhances the completeness of protein sequence clustering while maintaining homogeneity to produce more biologically representative clusters.
Key Features:
- Balancing Homogeneity and Completeness: Complet_plus merges closely related clusters based on verified structural relationships defined by the SCOPe classification to improve completeness while preserving homogeneity.
- Scalability: The method demonstrates linear scalability in runtime relative to the number of clusters and is applicable to large datasets such as a COG dataset containing over 3 million sequences.
- Integration with Existing Tools: Complet_plus functions as a post-processing step for clustering outputs from MMseqs2's clusterupdate, CD-HIT, and linclust.
- Quantitative Improvements: Application with MMseqs2's clusterupdate increased V-measure by 0.09 at the SCOPe superfamily level and 0.05 at the family level and produced substantial increases in Adjusted Mutual Information (AMI) and Adjusted Rand Index (ARI) against biological classifications.
Scientific Applications:
- Proteome-scale clustering: Improves biological representativeness of clusters in large-scale proteomic datasets.
- Structural classification benchmarking: Enhances mapping between sequence clusters and structural classifications such as SCOPe for validating clustering accuracy.
- Functional and evolutionary inference: Supports more accurate inference of protein functions and evolutionary relationships from clustered sequences.
Methodology:
Complet_plus operates as a post-processing step that merges clusters based on verified structural relationships (SCOPe) and demonstrates linear runtime scaling relative to the number of clusters.
Topics
Details
- License:
- Not licensed
- Tool Type:
- command-line tool
- Programming Languages:
- Shell, Python
- Added:
- 3/9/2023
- Last Updated:
- 11/24/2024
Operations
Publications
Nguyen R, Sokhansanj BA, Polikar R, Rosen GL. Complet+: a computationally scalable method to improve completeness of large-scale protein sequence clustering. PeerJ. 2023;11:e14779. doi:10.7717/peerj.14779. PMID:36785708. PMCID:PMC9921987.
DOI: 10.7717/peerj.14779
PMID: 36785708
PMCID: PMC9921987
Funding: - NSF grants: #1919691, #1936782, #1936791, #2107108