CircSI-SSL

CircSI-SSL applies self-supervised learning to identify circular RNA (circRNA) protein-binding sites from sequence-derived representations.


Key Features:

  • Self-Supervised Learning: Employs a self-supervised learning paradigm to learn from unlabeled circRNA sequence data.
  • Multi-View Feature Coding and RNA_Transformer: Uses multiple feature coding schemes combined with an RNA_Transformer model for cross-view sequence prediction to learn mutual information across multi-view representations.
  • Fine-Tuning with Limited Labels: Fine-tunes the pretrained model using a minimal number of labeled samples to improve generalization from scarce annotations.
  • Empirical Performance on circRNA Datasets: Validated on six circRNA datasets and reported to outperform existing algorithms, including at a low training:test ratio (1:9).
  • Transplantation to linRNA Datasets: Maintains performance on six linear RNA (linRNA) datasets without modifications to network architecture or hyperparameters, indicating scalability.

Scientific Applications:

  • Regulatory role analysis: Predicts circRNA protein-binding sites to facilitate study of circRNA regulatory roles in gene expression.
  • Disease mechanism and therapeutic target investigation: Supports exploration of circRNA involvement in various diseases and identification of potential therapeutic targets.
  • Biomarker discovery: Aids discovery of circRNA-related biomarkers through predicted circRNA–protein interaction sites.

Methodology:

Pretrain an RNA_Transformer via self-supervised cross-view sequence prediction using multiple feature coding schemes to learn mutual information across representations, then fine-tune the model with a minimal number of labeled samples.

Topics

Details

License:
MIT
Cost:
Free of charge
Tool Type:
command-line tool
Programming Languages:
Python
Added:
5/18/2024
Last Updated:
11/24/2024

Operations

Data Inputs & Outputs

Publications

Cao C, Wang C, Yang S, Zou Q. CircSI-SSL: circRNA-binding site identification based on self-supervised learning. Bioinformatics. 2024;40(1). doi:10.1093/bioinformatics/btae004. PMID:38180876. PMCID:PMC10789309.

PMID: 38180876
Funding: - National Natural Science Foundation of China: 62231013, 62250028, 62271329 - Sichuan Provincial Science Fund for Distinguished Young Scholars: 2021JDJQ0025 - Shenzhen Polytechnic: 6022310036K, 6023310037K - Municipal Government of Quzhou: 2022D040