CUBAP

CUBAP analyzes codon usage biases across human populations to quantify population-specific differences in codon-related metrics and relate them to potential functional and disease-associated effects.


Key Features:

  • 1000 Genomes Project data: Uses variant and sequence data from the 1000 Genomes Project to assess population variation.
  • Gene coverage: Evaluates codon-related metrics across 17,634 genes.
  • Codon frequencies: Computes codon usage frequencies per gene and population.
  • Codon aversion: Quantifies absence or underrepresentation of specific codons across populations.
  • Identical codon pairing: Detects and compares occurrences of adjacent identical codons between populations.
  • Co-tRNA codon pairing: Assesses pairing of codons decoded by the same tRNA (co-tRNA) across genes and populations.
  • Ramp sequences: Analyzes ramp sequence patterns per gene and population.
  • Nucleotide composition: Measures nucleotide composition across genes and populations.
  • Population variation in codon pairing: Identifies significant codon-pairing variation between populations for 35.8% of genes.
  • Ancestry prediction: Predicts individual place of origin with reported accuracies of 98.8% for African individuals and 100% for East Asian individuals.
  • IRGM CTG pairing bias: Detects decreased CTG pairing in the IRGM gene in East Asian and African populations and notes correlation with a reduced association of rs10065172 with Crohn's disease.

Scientific Applications:

  • Genome-wide association studies: Supports interpretation of synonymous variants and codon-level biases in GWAS contexts.
  • Candidate-gene analysis: Enables gene-specific and codon-specific analyses to evaluate candidate genes.
  • Functional interpretation of synonymous variants: Aids prediction of population-specific impacts of synonymous variants.
  • Population genetics and ancestry inference: Uses codon usage patterns to categorize genetic biases unique to populations and infer place of origin.
  • Disease-association analysis: Facilitates investigation of codon usage changes in disease contexts, exemplified by IRGM CTG pairing and rs10065172 in Crohn's disease.

Methodology:

Uses 1000 Genomes Project data to calculate codon frequencies, codon aversion, identical codon pairing, co-tRNA codon pairing, ramp sequences, and nucleotide composition across 17,634 genes and assesses population-specific differences and predictive assignments.

Topics

Details

Tool Type:
web application
Programming Languages:
Python
Added:
1/18/2021
Last Updated:
2/18/2021

Operations

Publications

Hodgman MW, Miller JB, Meurs TE, Kauwe JSK. CUBAP: an interactive web portal for analyzing codon usage biases across populations. Nucleic Acids Research. 2020;48(19):11030-11039. doi:10.1093/nar/gkaa863. PMID:33045750. PMCID:PMC7641757.

PMID: 33045750
PMCID: PMC7641757
Funding: - National Institute on Aging: RF1 AG054052

Documentation

Links