GeneMark-EP
GeneMark-EP integrates external protein evidence via the ProtHint pipeline to improve ab initio gene prediction and model parameter estimation in eukaryotic genomes.
Key Features:
- Integration with Protein Databases: Incorporates an external protein database as evidence by mapping proteins to the genome to inform gene prediction.
- ProtHint Pipeline: Uses ProtHint for large-scale protein-to-genome mapping and for extracting hints about splice sites and translation start and stop sites from protein alignments.
- Self-Training Mechanism: Employs iterative self-training that refines gene prediction model parameters using protein-derived hints.
- GeneMark-EP+ Coordinate Adjustment: In GeneMark-EP+ mode, adjusts predicted gene coordinates when they conflict with high-confidence protein-derived hints.
- Relation to GeneMark-ES and GeneMark-ET: Builds on the methodologies and self-training frameworks established by GeneMark-ES and GeneMark-ET.
Scientific Applications:
- Large eukaryotic genome annotation: Enhances gene prediction accuracy in large and complex eukaryotic genomes by incorporating protein evidence.
- Gene structure refinement: Improves identification of splice sites and translation initiation/termination sites to produce more accurate gene coordinates and structures.
Methodology:
Proteins are mapped to the genome with the ProtHint pipeline; mapping extracts hints on splice sites and translation initiation/termination; those hints are used to refine model parameter estimation via self-training; in GeneMark-EP+ predicted gene coordinates are adjusted when they conflict with reliable protein-derived hints.
Topics
Details
- Tool Type:
- web application
- Programming Languages:
- Shell, Perl, Python
- Added:
- 1/18/2021
- Last Updated:
- 1/22/2021
Operations
Publications
Brůna T, Lomsadze A, Borodovsky M. GeneMark-EP and -EP+: eukaryotic gene prediction with self-training in the space of genes and proteins. Unknown Journal. 2020. doi:10.1101/2019.12.31.891218.