FragGeneScan

FragGeneScan predicts protein-coding regions from short and error-prone sequencing reads, with emphasis on metagenomic samples.


Key Features:

  • Hidden Markov Model (HMM): Uses an HMM framework that integrates sequencing error models and codon usage patterns for gene prediction.
  • Error Tolerance: Maintains high prediction accuracy in the presence of sequencing errors and shows improved performance over MetaGene as error rates increase.
  • Short Read Optimization: Optimized for short reads and outperforms Glimmer and MetaGene on fragmented metagenomic sequences.
  • Enhanced Accuracy: For 400-base reads with a 1% sequencing error rate it improves prediction accuracy by approximately 62%, and for 100-base error-free reads it enhances accuracy by about 18%.
  • Gene Recovery: Recovers a substantially higher number of genes compared to MetaGene and identifies over 90% of genes detected through homology searches.

Scientific Applications:

  • Metagenomic gene prediction: Predicts protein-coding regions directly from short reads without relying on genome assembly, enabling analysis of complex environmental samples.
  • Gene recovery and discovery: Recovers more genes than MetaGene and uncovers novel genes with no known homologs in protein databases, supporting studies of microbial diversity and function.

Methodology:

Integrates sequencing error models and codon usage statistics within a hidden Markov model framework.

Topics

Collections

Details

License:
GPL-3.0
Maturity:
Mature
Tool Type:
command-line tool
Operating Systems:
Linux
Programming Languages:
Perl, C
Added:
1/13/2017
Last Updated:
11/24/2024

Operations

Data Inputs & Outputs

Gene prediction

Publications

Rho M, Tang H, Ye Y. FragGeneScan: predicting genes in short and error-prone reads. Nucleic Acids Research. 2010;38(20):e191-e191. doi:10.1093/nar/gkq747. PMID:20805240. PMCID:PMC2978382.

Mareuil F, Doppelt-Azeroual O, Ménager H. A public Galaxy platform at Pasteur used as an execution engine for web services. Unknown Journal. 2017. doi:10.7490/f1000research.1114334.1.

Documentation

Links