SAdLSA
SAdLSA generates protein sequence alignments by using deep learning trained on structural alignments to detect structural relationships and decode aspects of the protein folding code for sequences with low identity.
Key Features:
- Deep learning on structural alignments: Trained on thousands of experimentally determined structural alignments to implicitly learn the protein folding code.
- Twilight-zone detection: Recognizes and aligns structurally related proteins even when sequence identity is low.
- Cross-fold generalization: Learns patterns that generalize across fold types (e.g., models trained on pure α-helical proteins can identify β-sheet domains).
- Comparative performance: Demonstrates approximately 150% improvement in generating pairwise alignments and about 50% greater effectiveness at identifying proteins with optimal alignments within a sequence library compared to HHsearch.
- Computational efficiency: GPU-accelerated implementation with O(N) time complexity for rapid processing of large datasets.
- Benchmarking: Trained and evaluated on diverse and challenging datasets of structural alignments.
Scientific Applications:
- Evolutionary analysis: Identification of distant homologs and structural relationships for evolutionary studies.
- Function annotation: Improved functional inference by detecting structural similarity beyond sequence identity.
- Structural prediction and comparison: Support for structure-aware alignment tasks relevant to structural prediction and domain comparison.
Methodology:
Deep learning models trained on datasets of experimentally determined structural alignments, benchmarking against HHsearch, and a GPU-accelerated implementation with O(N) time complexity.
Topics
Details
- Added:
- 1/18/2021
- Last Updated:
- 11/24/2024
Operations
Publications
Gao M, Skolnick J. A novel sequence alignment algorithm based on deep learning of the protein folding code. Bioinformatics. 2020;37(4):490-496. doi:10.1093/bioinformatics/btaa810. PMID:32960943. PMCID:PMC8599902.