ingap-cdg
ingap-cdg constructs full-length, non-redundant coding sequences (CDSs) directly from unassembled transcriptome reads to enable accurate gene prediction without closely related reference genomes.
Key Features:
- Codon-Based De Bruijn Graph: Employs a codon-based de Bruijn graph to assemble coding sequences from fragmented transcriptomic data.
- Machine Learning Integration: Uses machine learning algorithms to filter false positives and distinguish true coding sequences from sequencing artifacts.
- Robustness to Sequencing Errors: Maintains prediction accuracy despite sequencing errors and variability in read lengths.
- Increased Predicted CDS Length: Produces longer predicted coding sequences compared to other existing methods, improving completeness of gene representations.
Scientific Applications:
- Reference-limited transcriptome analysis: Enables CDS prediction in studies lacking closely related reference genomes, including non-model organisms and newly discovered species.
- Gene structure and function inference from raw reads: Facilitates reconstruction of coding sequences directly from raw transcriptome reads to support downstream gene structure and functional analyses.
Methodology:
Assembly simplification using a codon-based de Bruijn graph to construct coding sequences from fragmented transcriptomic data, and false-positive filtering using machine learning algorithms to distinguish true CDSs from artifacts.
Topics
Details
- License:
- CC-BY-4.0
- Tool Type:
- command-line tool
- Operating Systems:
- Linux, Mac
- Added:
- 8/12/2018
- Last Updated:
- 11/25/2024
Operations
Publications
Peng G, Ji P, Zhao F. A novel codon-based de Bruijn graph algorithm for gene construction from unassembled transcriptomes. Genome Biology. 2016;17(1). doi:10.1186/s13059-016-1094-x. PMID:27855707. PMCID:PMC5114782.
PMID: 27855707
PMCID: PMC5114782
Funding: - National Natural Science Foundation of China: 31671364, 91131013
- Strategic Priority Research Program of the Chinese Academy of Sciences: XDB13000000