protea
protea identifies protein-coding sequences in genomic DNA by detecting coding-specific substitution patterns and enforcing reading-frame consistency without requiring sequence alignments, for comparative genomics–based gene prediction.
Key Features:
- Alignment-free detection: Operates without requiring sequence alignments, enabling analysis of diverse sets of DNA sequences.
- Substitution pattern analysis: Leverages substitution patterns characteristic of protein-coding sequences to identify coding signals.
- Reading-frame consistency: Maintains and evaluates consistency within reading frames to distinguish coding from non-coding regions.
- Comparative genomics conservation detection: Identifies evolutionarily conserved protein-coding regions across genomes.
- Coding versus non-coding discrimination: Differentiates coding regions from non-coding regions based on substitution patterns and frame consistency.
- Complementary to homology-based methods: Serves as a complement to homology search and statistical gene prediction approaches by detecting conserved coding regions that may be overlooked.
Scientific Applications:
- Gene prediction: Identification of protein-coding genes in genomic DNA sequences.
- Comparative genomics: Detection of conserved coding regions across multiple genomes.
- Evolutionary and functional analysis: Uncovering conserved coding regions to inform functional and evolutionary insights.
- Annotation complementarity: Complementing homology-based annotation and statistical gene prediction workflows.
Methodology:
Uses a comparative genomics approach to detect evolutionarily conserved protein-coding regions by analyzing coding-specific substitution patterns and enforcing reading-frame consistency in an alignment-free manner.
Topics
Details
- Tool Type:
- desktop application, web application
- Operating Systems:
- Linux, Windows, Mac
- Added:
- 8/3/2017
- Last Updated:
- 11/25/2024
Operations
Publications
Fontaine A, Touzet H. Computational identification of protein-coding sequences by comparative analysis. International Journal of Data Mining and Bioinformatics. 2009;3(2):160. doi:10.1504/ijdmb.2009.024849. PMID:19517987.
PMID: 19517987