EasyGene

EasyGene predicts protein-coding genes in prokaryotic genomes using a hidden Markov model (HMM) and Swiss-Prot–derived training sets to assign statistical significance to open reading frames (ORFs) and standardize annotations.


Key Features:

  • Automated Gene Prediction: Identifies genes by analyzing open reading frames (ORFs) and generating a list of putative genes with associated likelihoods.
  • Statistical Significance Assessment: Uses a hidden Markov model (HMM) to evaluate the probability of an ORF being a genuine gene, considering ORF score and length.
  • High-Quality Training Set Utilization: Automatically extracts a training set of genes from the genome using similarities in Swiss-Prot to estimate HMM parameters.
  • Performance Evaluation and Comparison: Has been tested on 143 prokaryotic genomes and identified discrepancies between existing annotations (GenBank, RefSeq) and its predictions, including over-annotation in GC-rich genomes and under-annotation.
  • Standardized Annotation Method: Provides a consistent annotation approach that facilitates transparent genome comparisons and has shown slight performance gains over existing annotations in evaluations.
  • Automated Pipeline: Implements an automated computational pipeline from raw genome input to generation of statistically scored lists of putative genes.

Scientific Applications:

  • Genome Annotation: Produces standardized gene annotations for prokaryotic genomes.
  • Comparative Genomics: Enables consistent cross-species or strain comparisons by applying uniform annotation criteria.
  • Error Correction: Detects annotation errors such as incorrect start codons, over-annotated ORFs, and missing genes.

Methodology:

Parameters are estimated by training a hidden Markov model (HMM) on genome-derived gene sets extracted via Swiss-Prot similarity; ORFs are scored by length and HMM-derived probabilities to produce lists of putative genes with statistical significance.

Topics

Details

Tool Type:
web application
Operating Systems:
Linux, Windows, Mac
Added:
7/27/2015
Last Updated:
11/25/2024

Operations

Data Inputs & Outputs

Publications

Nielsen P, Krogh A. Large-scale prokaryotic gene prediction and comparison to genome annotation. Bioinformatics. 2005;21(24):4322-4329. doi:10.1093/bioinformatics/bti701. PMID:16249266.

Larsen TS, Krogh A. EasyGene – a prokaryotic gene finder that ranks ORFs by statistical significance. BMC Bioinformatics. 2003;4(1). doi:10.1186/1471-2105-4-21. PMID:12783628. PMCID:PMC521197.

Documentation