CADD

CADD predicts the deleteriousness of human genetic variants by integrating diverse genomic annotations into a unified C score to prioritize functional, deleterious, and pathogenic SNVs, insertions, and deletions.


Key Features:

  • Integration of annotations: Combines over 60 genomic features including conservation metrics, functional annotations, pathogenicity signals, regulatory effects, and associations with complex traits into a single score.
  • Machine learning classifier: Uses a support vector machine (SVM) trained to distinguish high-frequency human-derived alleles from simulated variants to prioritize functional and deleterious variants.
  • Precomputed SNV coverage: Provides precomputed scores for all possible ~8.6 billion human single-nucleotide variants and supports scoring of short insertions and deletions.
  • Biological correlation: C scores correlate with allelic diversity, functional annotations, pathogenicity, disease severity, experimentally measured regulatory effects, and GWAS associations.
  • Splicing prediction (cadd_phred): The cadd_phred extension incorporates deep neural network (DNN) splicing scores to predict splicing effects beyond canonical donor and acceptor dinucleotides.
  • Genome build support: Version 1.4 includes support for the human genome build GRCh38.

Scientific Applications:

  • Mendelian disorder variant prioritization: Prioritizes candidate causal variants in severe Mendelian disease analyses.
  • GWAS interpretation: Ranks and prioritizes variants underlying genome-wide association study signals using integrated annotations.
  • Exome and sequencing analyses: Supports variant effect prediction and prioritization in exome and genome sequencing workflows, including short indels.
  • Splicing impact detection: Identifies variants likely to affect splicing, including those outside canonical splice-site dinucleotides.

Methodology:

Integrates >60 genomic annotations, uses an SVM trained on high-frequency human-derived alleles versus simulated variants, provides precomputed scores for ~8.6 billion SNVs and scores for short indels, and incorporates DNN-derived splicing scores in the cadd_phred extension; version 1.4 adds GRCh38 support.

Topics

Collections

Details

License:
Unlicense
Maturity:
Mature
Cost:
Free of charge
Tool Type:
web application
Operating Systems:
Linux, Windows, Mac
Added:
7/4/2017
Last Updated:
11/24/2024

Operations

Publications

Kircher M, Witten DM, Jain P, O'Roak BJ, Cooper GM, Shendure J. A general framework for estimating the relative pathogenicity of human genetic variants. Nature Genetics. 2014;46(3):310-315. doi:10.1038/ng.2892. PMID:24487276. PMCID:PMC3992975.

Rentzsch P, Witten D, Cooper GM, Shendure J, Kircher M. CADD: predicting the deleteriousness of variants throughout the human genome. Nucleic Acids Research. 2018;47(D1):D886-D894. doi:10.1093/nar/gky1016. PMID:30371827. PMCID:PMC6323892.

PMID: 30371827
PMCID: PMC6323892
Funding: - National Cancer Institute: 1R01CA197139 - National Human Genome Research Institute: 1U54HG006493

Rentzsch P, Schubach M, Shendure J, Kircher M. CADD-Splice—improving genome-wide variant effect prediction using deep learning-derived splice scores. Genome Medicine. 2021;13(1). doi:10.1186/s13073-021-00835-9. PMID:33618777. PMCID:PMC7901104.

PMID: 33618777
PMCID: PMC7901104
Funding: - National Cancer Institute: 1R01CA197139

Documentation

Downloads