LncRNA-ID

LncRNA-ID evaluates the coding potential of transcripts, focusing on long non-coding RNAs (lncRNAs >200 nucleotides), by integrating sequence, ribosome profiling, and conservation features for transcript classification.


Key Features:

  • Random forest algorithm: Implements a random forest classifier to predict transcript coding potential.
  • Sequence characteristics of putative open reading frames (ORFs): Analyzes ORF sequence features to detect signals of protein-coding capacity.
  • Translation scores from ribosomal coverage: Computes translation scores based on ribosome profiling (ribosomal coverage) to capture evidence of translation.
  • Conservation against characterized protein families: Assesses evolutionary conservation by comparing sequences to characterized protein families to infer similarity to known proteins.

Scientific Applications:

  • lncRNA annotation across species: Annotates lncRNAs across various species by distinguishing them from protein-coding transcripts.
  • Transcript classification: Classifies transcripts by coding potential to separate lncRNAs from protein-coding genes.
  • Support for functional studies: Provides classification evidence to inform studies of lncRNA functions and their regulatory roles in gene expression.

Methodology:

Applies a random forest machine-learning model that integrates ORF sequence-derived features, ribosome profiling-based translation scores, and conservation comparisons against characterized protein families.

Topics

Details

Tool Type:
command-line tool
Operating Systems:
Linux
Programming Languages:
C++
Added:
8/3/2017
Last Updated:
11/25/2024

Operations

Publications

Achawanantakun R, Chen J, Sun Y, Zhang Y. LncRNA-ID: Long non-coding RNA IDentification using balanced random forests. Bioinformatics. 2015;31(24):3897-3905. doi:10.1093/bioinformatics/btv480. PMID:26315901.

Documentation

Links