CUBAP
CUBAP analyzes codon usage biases across human populations to quantify population-specific differences in codon-related metrics and relate them to potential functional and disease-associated effects.
Key Features:
- 1000 Genomes Project data: Uses variant and sequence data from the 1000 Genomes Project to assess population variation.
- Gene coverage: Evaluates codon-related metrics across 17,634 genes.
- Codon frequencies: Computes codon usage frequencies per gene and population.
- Codon aversion: Quantifies absence or underrepresentation of specific codons across populations.
- Identical codon pairing: Detects and compares occurrences of adjacent identical codons between populations.
- Co-tRNA codon pairing: Assesses pairing of codons decoded by the same tRNA (co-tRNA) across genes and populations.
- Ramp sequences: Analyzes ramp sequence patterns per gene and population.
- Nucleotide composition: Measures nucleotide composition across genes and populations.
- Population variation in codon pairing: Identifies significant codon-pairing variation between populations for 35.8% of genes.
- Ancestry prediction: Predicts individual place of origin with reported accuracies of 98.8% for African individuals and 100% for East Asian individuals.
- IRGM CTG pairing bias: Detects decreased CTG pairing in the IRGM gene in East Asian and African populations and notes correlation with a reduced association of rs10065172 with Crohn's disease.
Scientific Applications:
- Genome-wide association studies: Supports interpretation of synonymous variants and codon-level biases in GWAS contexts.
- Candidate-gene analysis: Enables gene-specific and codon-specific analyses to evaluate candidate genes.
- Functional interpretation of synonymous variants: Aids prediction of population-specific impacts of synonymous variants.
- Population genetics and ancestry inference: Uses codon usage patterns to categorize genetic biases unique to populations and infer place of origin.
- Disease-association analysis: Facilitates investigation of codon usage changes in disease contexts, exemplified by IRGM CTG pairing and rs10065172 in Crohn's disease.
Methodology:
Uses 1000 Genomes Project data to calculate codon frequencies, codon aversion, identical codon pairing, co-tRNA codon pairing, ramp sequences, and nucleotide composition across 17,634 genes and assesses population-specific differences and predictive assignments.
Topics
Details
- Tool Type:
- web application
- Programming Languages:
- Python
- Added:
- 1/18/2021
- Last Updated:
- 2/18/2021
Operations
Publications
Hodgman MW, Miller JB, Meurs TE, Kauwe JSK. CUBAP: an interactive web portal for analyzing codon usage biases across populations. Nucleic Acids Research. 2020;48(19):11030-11039. doi:10.1093/nar/gkaa863. PMID:33045750. PMCID:PMC7641757.
DOI: 10.1093/NAR/GKAA863
PMID: 33045750
PMCID: PMC7641757
Funding: - National Institute on Aging: RF1 AG054052
Documentation
User manual
https://cubap.readthedocs.io/Links
Repository
https://github.com/kauwelab/cubap