Seq-SetNet

Seq-SetNet leverages multiple sequence alignments (MSAs) to predict residue-level structural properties such as secondary structure and backbone torsion angles.


Key Features:

  • Encoding Module: Processes each component homologue in an MSA individually and encodes residue-specific mutations, insertions, deletions, and long-distance residue correlations into context-specific features for every residue.
  • Aggregation Module: Consolidates encoded features from all homologues and transforms the aggregated representation into residue-level structural properties for the query protein.
  • Automatic Feature Learning: Learns effective features directly from MSAs within the encoding module, removing the need for manual feature engineering.
  • Symmetric Aggregation Functions: Uses order-invariant (symmetric) aggregation functions to ensure predictions are independent of the sequence order of homologues in the MSA.

Scientific Applications:

  • Residue-level structural prediction: Predicts secondary structure and backbone torsion angles of protein residues directly from MSAs.
  • Protein folding, function, and interaction studies: Provides residue-resolved structural information to support analyses of protein folding, molecular function, and protein–protein or protein–ligand interactions.

Methodology:

Applies an explicit "encoding and aggregation" strategy that individually encodes each homologue in an MSA, aggregates those encoded features with symmetric functions, and transforms the aggregated features into residue-level structural properties, with feature representations learned automatically within the encoding module.

Topics

Details

License:
MIT
Cost:
Free of charge
Tool Type:
command-line tool
Operating Systems:
Mac, Linux, Windows
Programming Languages:
Python
Added:
5/17/2022
Last Updated:
5/17/2022

Operations

Publications

Ju F, Zhu J, Zhang Q, Wei G, Sun S, Zheng W, Bu D. Seq-SetNet: directly exploiting multiple sequence alignment for protein secondary structure prediction. Bioinformatics. 2021;38(4):990-996. doi:10.1093/bioinformatics/btab777. PMID:34849579.

PMID: 34849579
Funding: - National Key Research and Development Program of China: 2020YFA0907000 - National Natural Science Foundation of China: 31770775, 62072435, 82130055