GIIRA
GIIRA predicts genes in prokaryotic and eukaryotic genomes by integrating RNA-Seq data and reallocating ambiguously mapped reads to improve identification of expressed coding regions.
Key Features:
- RNA-Seq integration: Uses high-throughput RNA-Seq data to inform gene prediction.
- Ambiguous read handling: Retains and explicitly handles ambiguously mapped reads rather than discarding them.
- Candidate region extraction: Extracts candidate genomic regions based on a sufficient number of RNA-Seq read mappings.
- Maximum-flow reassignment: Employs a maximum-flow algorithm to reassign ambiguous reads to their most probable origin.
- Cross-domain support: Applicable to both prokaryotic and eukaryotic genomes for coding region identification.
- Validation: Evaluated on simulated and real datasets and reported improved performance compared to existing methods that incorporate RNA-Seq.
- Implementation: Implemented in Java.
Scientific Applications:
- Gene prediction: Identification of expressed coding regions from RNA-Seq data in prokaryotic and eukaryotic genomes.
- Recovery of genes with ambiguous support: Rescue of genes predominantly supported by ambiguously mapped reads.
- Method benchmarking: Comparative evaluation of gene-finding accuracy on simulated and real RNA-Seq datasets.
Methodology:
Extracts candidate regions based on a threshold of RNA-Seq read mappings and employs a maximum-flow algorithm to reassign ambiguously mapped reads to their most probable origin; implemented in Java.
Topics
Details
- Tool Type:
- command-line tool
- Operating Systems:
- Linux, Windows, Mac
- Programming Languages:
- Java
- Added:
- 8/3/2017
- Last Updated:
- 11/25/2024
Operations
Publications
Zickmann F, Lindner MS, Renard BY. GIIRA—RNA-Seq driven gene finding incorporating ambiguous reads. Bioinformatics. 2013;30(5):606-613. doi:10.1093/bioinformatics/btt577. PMID:24123675.
PMID: 24123675