CEGMA

CEGMA identifies and maps a conserved set of core eukaryotic genes (CEGs) to annotate exon–intron structures in draft eukaryotic genome assemblies and to assess completeness of gene catalogs.


Key Features:

  • Core gene set: defines a set of highly conserved protein families (core eukaryotic genes, CEGs) prevalent across diverse eukaryotic organisms.
  • Low-copy markers: targets CEGs characterized by extreme conservation and low copy numbers in higher eukaryotes to serve as reliable single-copy markers.
  • Profile HMMs: employs profile-hidden Markov models (HMMs) to detect conserved protein families in genomic sequences.
  • Exon–intron mapping: uses a mapping procedure to identify exon–intron structures of conserved protein families within novel genomic sequences.
  • Initial gene annotations: constructs an initial set of reliable gene annotations in draft-stage genomes.
  • Assembly completeness metric: reports the proportion of CEGs mapped within a draft genome to evaluate completeness of gene catalogs.
  • Gene-finder training and validation: provides annotations useful for training gene-finding algorithms and validating automatic predictions.
  • Complementary metrics: complements N50 length and x-fold coverage by directly assessing gene space in assemblies.

Scientific Applications:

  • Assembly quality assessment: assessing the completeness of gene catalogs in draft eukaryotic genome assemblies.
  • Gene-prediction development: training and validating gene-finding algorithms using conserved gene annotations.
  • Comparative assembly metrics: providing a gene-space-based complement to N50 length and x-fold coverage for sequencing and assembly evaluation.
  • Initial annotation in low-data projects: generating initial reliable gene annotations when experimental evidence is limited or absent.

Methodology:

Defines a set of highly conserved protein families (CEGs), builds and applies profile-hidden Markov models (HMMs) to map those families to genomic sequences, uses a mapping procedure to identify exon–intron structures, and reports the proportion of CEGs mapped as an assembly completeness metric.

Topics

Details

Tool Type:
command-line tool
Operating Systems:
Linux, Mac
Programming Languages:
Perl
Added:
12/18/2017
Last Updated:
11/25/2024

Operations

Data Inputs & Outputs

Publications

Parra G, Bradnam K, Ning Z, Keane T, Korf I. Assessing the gene space in draft genomes. Nucleic Acids Research. 2008;37(1):289-297. doi:10.1093/nar/gkn916. PMID:19042974. PMCID:PMC2615622.

Parra G, Bradnam K, Korf I. CEGMA: a pipeline to accurately annotate core genes in eukaryotic genomes. Bioinformatics. 2007;23(9):1061-1067. doi:10.1093/bioinformatics/btm071. PMID:17332020.

Documentation

Links