LncRNA-ID
LncRNA-ID evaluates the coding potential of transcripts, focusing on long non-coding RNAs (lncRNAs >200 nucleotides), by integrating sequence, ribosome profiling, and conservation features for transcript classification.
Key Features:
- Random forest algorithm: Implements a random forest classifier to predict transcript coding potential.
- Sequence characteristics of putative open reading frames (ORFs): Analyzes ORF sequence features to detect signals of protein-coding capacity.
- Translation scores from ribosomal coverage: Computes translation scores based on ribosome profiling (ribosomal coverage) to capture evidence of translation.
- Conservation against characterized protein families: Assesses evolutionary conservation by comparing sequences to characterized protein families to infer similarity to known proteins.
Scientific Applications:
- lncRNA annotation across species: Annotates lncRNAs across various species by distinguishing them from protein-coding transcripts.
- Transcript classification: Classifies transcripts by coding potential to separate lncRNAs from protein-coding genes.
- Support for functional studies: Provides classification evidence to inform studies of lncRNA functions and their regulatory roles in gene expression.
Methodology:
Applies a random forest machine-learning model that integrates ORF sequence-derived features, ribosome profiling-based translation scores, and conservation comparisons against characterized protein families.
Topics
Details
- Tool Type:
- command-line tool
- Operating Systems:
- Linux
- Programming Languages:
- C++
- Added:
- 8/3/2017
- Last Updated:
- 11/25/2024
Operations
Publications
Achawanantakun R, Chen J, Sun Y, Zhang Y. LncRNA-ID: Long non-coding RNA IDentification using balanced random forests. Bioinformatics. 2015;31(24):3897-3905. doi:10.1093/bioinformatics/btv480. PMID:26315901.
PMID: 26315901