CircSI-SSL
CircSI-SSL applies self-supervised learning to identify circular RNA (circRNA) protein-binding sites from sequence-derived representations.
Key Features:
- Self-Supervised Learning: Employs a self-supervised learning paradigm to learn from unlabeled circRNA sequence data.
- Multi-View Feature Coding and RNA_Transformer: Uses multiple feature coding schemes combined with an RNA_Transformer model for cross-view sequence prediction to learn mutual information across multi-view representations.
- Fine-Tuning with Limited Labels: Fine-tunes the pretrained model using a minimal number of labeled samples to improve generalization from scarce annotations.
- Empirical Performance on circRNA Datasets: Validated on six circRNA datasets and reported to outperform existing algorithms, including at a low training:test ratio (1:9).
- Transplantation to linRNA Datasets: Maintains performance on six linear RNA (linRNA) datasets without modifications to network architecture or hyperparameters, indicating scalability.
Scientific Applications:
- Regulatory role analysis: Predicts circRNA protein-binding sites to facilitate study of circRNA regulatory roles in gene expression.
- Disease mechanism and therapeutic target investigation: Supports exploration of circRNA involvement in various diseases and identification of potential therapeutic targets.
- Biomarker discovery: Aids discovery of circRNA-related biomarkers through predicted circRNA–protein interaction sites.
Methodology:
Pretrain an RNA_Transformer via self-supervised cross-view sequence prediction using multiple feature coding schemes to learn mutual information across representations, then fine-tune the model with a minimal number of labeled samples.
Topics
Details
- License:
- MIT
- Cost:
- Free of charge
- Tool Type:
- command-line tool
- Programming Languages:
- Python
- Added:
- 5/18/2024
- Last Updated:
- 11/24/2024
Operations
Data Inputs & Outputs
Aggregation
Outputs
Publications
Cao C, Wang C, Yang S, Zou Q. CircSI-SSL: circRNA-binding site identification based on self-supervised learning. Bioinformatics. 2024;40(1). doi:10.1093/bioinformatics/btae004. PMID:38180876. PMCID:PMC10789309.
PMID: 38180876
PMCID: PMC10789309
Funding: - National Natural Science Foundation of China: 62231013, 62250028, 62271329
- Sichuan Provincial Science Fund for Distinguished Young Scholars: 2021JDJQ0025
- Shenzhen Polytechnic: 6022310036K, 6023310037K
- Municipal Government of Quzhou: 2022D040