Seq-SetNet
Seq-SetNet leverages multiple sequence alignments (MSAs) to predict residue-level structural properties such as secondary structure and backbone torsion angles.
Key Features:
- Encoding Module: Processes each component homologue in an MSA individually and encodes residue-specific mutations, insertions, deletions, and long-distance residue correlations into context-specific features for every residue.
- Aggregation Module: Consolidates encoded features from all homologues and transforms the aggregated representation into residue-level structural properties for the query protein.
- Automatic Feature Learning: Learns effective features directly from MSAs within the encoding module, removing the need for manual feature engineering.
- Symmetric Aggregation Functions: Uses order-invariant (symmetric) aggregation functions to ensure predictions are independent of the sequence order of homologues in the MSA.
Scientific Applications:
- Residue-level structural prediction: Predicts secondary structure and backbone torsion angles of protein residues directly from MSAs.
- Protein folding, function, and interaction studies: Provides residue-resolved structural information to support analyses of protein folding, molecular function, and protein–protein or protein–ligand interactions.
Methodology:
Applies an explicit "encoding and aggregation" strategy that individually encodes each homologue in an MSA, aggregates those encoded features with symmetric functions, and transforms the aggregated features into residue-level structural properties, with feature representations learned automatically within the encoding module.
Topics
Details
- License:
- MIT
- Cost:
- Free of charge
- Tool Type:
- command-line tool
- Operating Systems:
- Mac, Linux, Windows
- Programming Languages:
- Python
- Added:
- 5/17/2022
- Last Updated:
- 5/17/2022
Operations
Publications
Ju F, Zhu J, Zhang Q, Wei G, Sun S, Zheng W, Bu D. Seq-SetNet: directly exploiting multiple sequence alignment for protein secondary structure prediction. Bioinformatics. 2021;38(4):990-996. doi:10.1093/bioinformatics/btab777. PMID:34849579.