Conserved domain database CDD
Conserved Domain Database (CDD) annotates protein sequences with conserved domain footprints and associated functional sites to elucidate sequence, structure, and function relationships.
Key Features:
- Manually curated domain models: Provides manually curated domain models informed by protein three-dimensional (3D) structures for detailed domain footprints and functional site annotation.
- Multiple sequence alignments: Represents conserved protein domains using multiple sequence alignments that capture evolutionary conservation.
- Integration with external databases: Integrates domain models from public collections such as Pfam and SMART alongside contributions from NCBI collaborators.
- Hierarchical organization: Organizes domain models hierarchically to reflect evolutionary relationships among domain families and superfamilies.
- Clustering into superfamilies: Clusters redundant and homologous domain models into superfamilies to provide both specific family assignments and broader superfamily designations.
- Pre-computed Entrez annotations: Supplies pre-computed domain annotations for proteins tracked in NCBI's Entrez system.
- Sequence search via CD-Search and RPS-BLAST: Supports CD-Search and batch submissions that employ reverse-position-specific BLAST (RPS-BLAST) to identify potential domain matches.
- Structure–conservation mapping: Maps residue coordinates to evolutionary conservation and 3D structural information, with mapping and inspection supported by the molecular structure viewer Cn3D.
- Curation of novel families: Curates and classifies novel domain families emerging from new protein structure determinations and addresses unannotated "protein dark matter."
Scientific Applications:
- Protein annotation: Assigns conserved domains and functional sites to protein sequences for annotation and gene product naming.
- Function prediction: Informs prediction of molecular function based on conserved domain footprints and annotated functional residues.
- Domain architecture analysis: Elucidates protein domain architecture and relationships among domain families and superfamilies.
- Structure–function interpretation: Links domain models to 3D structural data to interpret structure–function relationships and map functional sites.
- Batch proteome annotation: Enables batch annotation of protein sets using CD-Search with RPS-BLAST for large-scale analyses.
Methodology:
CDD builds domain models from multiple sequence alignments informed by protein 3D structures, clusters redundant and homologous models into superfamilies and organizes them hierarchically, provides pre-computed annotations for Entrez proteins, and identifies matches via CD-Search using reverse-position-specific BLAST (RPS-BLAST) with residue-to-structure/conservation mapping supported by Cn3D.
Topics
Collections
Details
- Tool Type:
- web application
- Operating Systems:
- Linux, Windows, Mac
- Added:
- 9/11/2015
- Last Updated:
- 11/24/2024
Operations
Publications
Marchler-Bauer A. CDD: a Conserved Domain Database for protein classification. Nucleic Acids Research. 2004;33(Database issue):D192-D196. doi:10.1093/nar/gki069. PMID:15608175. PMCID:PMC540023.
Marchler-Bauer A, Zheng C, Chitsaz F, Derbyshire MK, Geer LY, Geer RC, Gonzales NR, Gwadz M, Hurwitz DI, Lanczycki CJ, Lu F, Lu S, Marchler GH, Song JS, Thanki N, Yamashita RA, Zhang D, Bryant SH. CDD: conserved domains and protein three-dimensional structure. Nucleic Acids Research. 2012;41(D1):D348-D352. doi:10.1093/nar/gks1243. PMID:23197659. PMCID:PMC3531192.
Wang J, Chitsaz F, Derbyshire MK, Gonzales NR, Gwadz M, Lu S, Marchler GH, Song JS, Thanki N, Yamashita RA, Yang M, Zhang D, Zheng C, Lanczycki CJ, Marchler-Bauer A. The conserved domain database in 2023. Nucleic Acids Research. 2022;51(D1):D384-D388. doi:10.1093/nar/gkac1096. PMID:36477806. PMCID:PMC9825596.
Marchler-Bauer A, Lu S, Anderson JB, Chitsaz F, Derbyshire MK, DeWeese-Scott C, Fong JH, Geer LY, Geer RC, Gonzales NR, Gwadz M, Hurwitz DI, Jackson JD, Ke Z, Lanczycki CJ, Lu F, Marchler GH, Mullokandov M, Omelchenko MV, Robertson CL, Song JS, Thanki N, Yamashita RA, Zhang D, Zhang N, Zheng C, Bryant SH. CDD: a Conserved Domain Database for the functional annotation of proteins. Nucleic Acids Research. 2010;39(Database):D225-D229. doi:10.1093/nar/gkq1189. PMID:21109532. PMCID:PMC3013737.
Marchler-Bauer A, Lu S, Anderson JB, Chitsaz F, Derbyshire MK, DeWeese-Scott C, Fong JH, Geer LY, Geer RC, Gonzales NR, Gwadz M, Hurwitz DI, Jackson JD, Ke Z, Lanczycki CJ, Lu F, Marchler GH, Mullokandov M, Omelchenko MV, Robertson CL, Song JS, Thanki N, Yamashita RA, Zhang D, Zhang N, Zheng C, Bryant SH. CDD: a Conserved Domain Database for the functional annotation of proteins. Nucleic Acids Research. 2010;39(Database):D225-D229. doi:10.1093/nar/gkq1189. PMID:21109532. PMCID:PMC3013737.
Lu S, Wang J, Chitsaz F, Derbyshire MK, Geer RC, Gonzales NR, Gwadz M, Hurwitz DI, Marchler GH, Song JS, Thanki N, Yamashita RA, Yang M, Zhang D, Zheng C, Lanczycki CJ, Marchler-Bauer A. CDD/SPARCLE: the conserved domain database in 2020. Nucleic Acids Research. 2019;48(D1):D265-D268. doi:10.1093/nar/gkz991. PMID:31777944. PMCID:PMC6943070.
Marchler-Bauer A, Anderson JB, Derbyshire MK, DeWeese-Scott C, Gonzales NR, Gwadz M, Hao L, He S, Hurwitz DI, Jackson JD, Ke Z, Krylov D, Lanczycki CJ, Liebert CA, Liu C, Lu F, Lu S, Marchler GH, Mullokandov M, Song JS, Thanki N, Yamashita RA, Yin JJ, Zhang D, Bryant SH. CDD: a conserved domain database for interactive domain family analysis. Nucleic Acids Research. 2007;35(Database):D237-D240. doi:10.1093/nar/gkl951. PMID:17135202. PMCID:PMC1751546.
Marchler-Bauer A. CDD: a database of conserved domain alignments with links to domain three-dimensional structure. Nucleic Acids Research. 2002;30(1):281-283. doi:10.1093/nar/30.1.281. PMID:11752315. PMCID:PMC99109.
Marchler-Bauer A, Anderson JB, Chitsaz F, Derbyshire MK, DeWeese-Scott C, Fong JH, Geer LY, Geer RC, Gonzales NR, Gwadz M, He S, Hurwitz DI, Jackson JD, Ke Z, Lanczycki CJ, Liebert CA, Liu C, Lu F, Lu S, Marchler GH, Mullokandov M, Song JS, Tasneem A, Thanki N, Yamashita RA, Zhang D, Zhang N, Bryant SH. CDD: specific functional annotation with the Conserved Domain Database. Nucleic Acids Research. 2009;37(Database):D205-D210. doi:10.1093/nar/gkn845. PMID:18984618. PMCID:PMC2686570.