CEGA

CEGA identifies conserved non-coding sequences (CNCs) and ultra-conserved elements (UCEs) across vertebrate species to map functionally constrained genomic regions using comparative genomics.


Key Features:

  • Conserved element detection: Identifies CNCs and UCEs that are under purifying selection across vertebrate lineages.
  • Clade-specific catalogs: Provides CNC sets for five vertebrate clades with counts of 24,488 (vertebrates), 241,575 (amniotes), 709,743 (Eutheria), 642,701 (Boreoeutheria), and 612,364 (Euarchontoglires), and total conserved spans ranging from ≈6 Mbp (vertebrates) to ≈119 Mbp (Euarchontoglires).
  • Comparative multi-species analysis: Uses multiple species comparisons to detect genomic elements preserved through evolutionary processes.
  • Phylogenetic modeling: Applies threshold-free phylogenetic modeling to assess conservation without arbitrary score cutoffs.
  • Global alignments of synteny blocks: Employs unbiased and sensitive global alignments of genomic synteny blocks for element detection.
  • Synteny detection via protein orthology: Identifies genomic synteny blocks using protein orthology relationships.
  • Function-agnostic identification: Detects conserved elements without requiring prior functional annotation.
  • Candidate regulatory regions: Reports conserved regions that can serve as candidates for enhancers, promoters, motifs, and other regulatory elements.

Scientific Applications:

  • Evolutionary genomics: Comparative analysis of functional constraint and heterogeneity in non-coding genomic sequences across vertebrate lineages.
  • Regulatory element discovery: Prioritization and selection of candidate enhancers, promoters, motifs, and other regulatory elements within conserved regions.
  • Phylogenetic conservation studies: Examination of conservation patterns at different phylogenetic depths (vertebrates, amniotes, Eutheria, Boreoeutheria, Euarchontoglires).
  • Comparative genomics: Cross-species identification of preserved genomic elements using synteny and orthology information.

Methodology:

Multiple species comparisons combined with threshold-free phylogenetic modeling applied to unbiased, sensitive global alignments of genomic synteny blocks that are identified using protein orthology.

Topics

Details

Tool Type:
web application
Operating Systems:
Linux, Windows, Mac
Added:
10/10/2016
Last Updated:
11/25/2024

Operations

Publications

Dousse A, Junier T, Zdobnov EM. CEGA—a catalog of conserved elements from genomic alignments. Nucleic Acids Research. 2015;44(D1):D96-D100. doi:10.1093/nar/gkv1163. PMID:26527719. PMCID:PMC4702837.