mkdom
mkdom identifies and clusters protein domains using recursive PSI-BLAST homology searches to delineate conserved protein domain building blocks across Bacteria, Archaea, and yeast.
Key Features:
- Domain Clustering: Clusters protein domains by assuming the shortest amino acid sequence represents a single domain and using that representation for querying and clustering.
- Recursive Homology Searches: Employs recursive PSI-BLAST searches to detect homologous sequences and delineate frequent domain building blocks.
- Cross-Domain Analysis: Identifies domains unique to specific domains of life and those shared among Bacteria, Archaea, and yeast, marking commonly shared domains as fundamental building blocks.
- Multi-Domain Protein Statistics: Computes statistics on multi-domain proteins and highlights organisms with unusually high numbers of large multi-domain proteins such as Bacillus subtilis and Mycobacterium tuberculosis.
- Domain Shuffling and Permutation Analysis: Detects examples of highly shuffled domains and circular permutations to inform studies of protein evolution and functional diversity.
- Genome-scale Analysis: Applied to analysis of protein domain shuffling across 17 completed microbial genomes.
Scientific Applications:
- Evolutionary Biology: Investigates conservation and diversification of protein domains across Bacteria, Archaea, and yeast.
- Structural Bioinformatics: Studies protein architecture and the modular nature of domains for structural interpretation.
- Comparative Genomics: Compares domain distribution and composition across multiple genomes to identify lineage-specific and shared domains.
- Protein Evolution Studies: Examines domain shuffling and circular permutation events to understand mechanisms of protein functional diversification.
Methodology:
Performs recursive PSI-BLAST homology searches and clusters domains using the shortest amino acid sequence as the representative single domain; applies these analyses across genomes (17 completed microbial genomes) to delineate domain building blocks, detect domain shuffling and circular permutations, and compute multi-domain protein statistics.
Topics
Details
- Tool Type:
- desktop application
- Operating Systems:
- Linux
- Programming Languages:
- Perl
- Added:
- 8/3/2017
- Last Updated:
- 11/25/2024
Operations
Publications
Gouzy J, Corpet F, Kahn D. Whole genome protein domain analysis using a new method for domain clustering. Computers & Chemistry. 1999;23(3-4):333-340. doi:10.1016/s0097-8485(99)00011-x. PMID:10404623.