EasyGene
EasyGene predicts protein-coding genes in prokaryotic genomes using a hidden Markov model (HMM) and Swiss-Prot–derived training sets to assign statistical significance to open reading frames (ORFs) and standardize annotations.
Key Features:
- Automated Gene Prediction: Identifies genes by analyzing open reading frames (ORFs) and generating a list of putative genes with associated likelihoods.
- Statistical Significance Assessment: Uses a hidden Markov model (HMM) to evaluate the probability of an ORF being a genuine gene, considering ORF score and length.
- High-Quality Training Set Utilization: Automatically extracts a training set of genes from the genome using similarities in Swiss-Prot to estimate HMM parameters.
- Performance Evaluation and Comparison: Has been tested on 143 prokaryotic genomes and identified discrepancies between existing annotations (GenBank, RefSeq) and its predictions, including over-annotation in GC-rich genomes and under-annotation.
- Standardized Annotation Method: Provides a consistent annotation approach that facilitates transparent genome comparisons and has shown slight performance gains over existing annotations in evaluations.
- Automated Pipeline: Implements an automated computational pipeline from raw genome input to generation of statistically scored lists of putative genes.
Scientific Applications:
- Genome Annotation: Produces standardized gene annotations for prokaryotic genomes.
- Comparative Genomics: Enables consistent cross-species or strain comparisons by applying uniform annotation criteria.
- Error Correction: Detects annotation errors such as incorrect start codons, over-annotated ORFs, and missing genes.
Methodology:
Parameters are estimated by training a hidden Markov model (HMM) on genome-derived gene sets extracted via Swiss-Prot similarity; ORFs are scored by length and HMM-derived probabilities to produce lists of putative genes with statistical significance.
Topics
Details
- Tool Type:
- web application
- Operating Systems:
- Linux, Windows, Mac
- Added:
- 7/27/2015
- Last Updated:
- 11/25/2024
Operations
Data Inputs & Outputs
Gene prediction
Publications
Nielsen P, Krogh A. Large-scale prokaryotic gene prediction and comparison to genome annotation. Bioinformatics. 2005;21(24):4322-4329. doi:10.1093/bioinformatics/bti701. PMID:16249266.
Larsen TS, Krogh A. EasyGene – a prokaryotic gene finder that ranks ORFs by statistical significance. BMC Bioinformatics. 2003;4(1). doi:10.1186/1471-2105-4-21. PMID:12783628. PMCID:PMC521197.