PIPENN
PIPENN predicts residue-level protein binding interfaces from protein sequence for protein-protein, protein-small-molecule, and protein-nucleotide (DNA/RNA) interactions.
Key Features:
- Residue-level prediction: Predicts binding interface residues at the amino-acid level using only protein sequence input.
- Interaction types: Covers protein-protein interactions (PPI), protein-small-molecule binding, and protein-nucleotide (DNA/RNA) binding.
- Deep learning architectures: Implements seven distinct deep learning architectures evaluated across datasets.
- Ensemble predictor: Combines multiple architectures into an ensemble that yields superior prediction accuracy compared to single architectures.
- Benchmarking sets: Evaluated across eleven independent test sets, including an independent test set denoted ZK448.
- Training dataset (BioDL): Uses the BioDL dataset comprising PPI data from the Protein Data Bank (PDB) and protein-ligand interactions from the BioLip database.
- Dataset augmentation: BioDL is augmented with curated homo- and heteromeric PPI datasets.
- Model development experiments: Assessed the impact of data features, spatial forms, encoding schemes, network initializations, loss functions, regularization mechanisms, and activation functions on predictor performance.
- Performance metrics: Benchmarked performance using area under the curve (AUC) metrics, reporting AUCs of 0.718 for PPI, 0.823 for protein-nucleotide, and 0.842 for protein-small-molecule interactions on ZK448.
Scientific Applications:
- Interface mapping: Identification of residue-level interfaces for studying signaling pathways, transcription, receptor activation, and enzymatic activity.
- Interaction characterization: Prediction of protein-protein, protein-small-molecule, and protein-nucleotide binding sites to support molecular interaction studies.
- Experimental prioritization: Prioritization of candidate interface residues for experimental validation or mutational analysis.
- Method benchmarking: Comparative evaluation and benchmarking of sequence-based protein interface predictors.
Methodology:
Seven distinct deep learning architectures were trained and evaluated, combined into an ensemble; evaluations used the BioDL dataset (PDB for PPI and BioLip for ligands) augmented with curated homo- and heteromeric PPI datasets, tested across eleven independent test sets including ZK448, with performance measured by AUC; development experiments varied data features, spatial forms, encoding schemes, network initializations, loss functions, regularization mechanisms, and activation functions.
Topics
Details
- License:
- GPL-3.0
- Cost:
- Free of charge
- Tool Type:
- command-line tool
- Operating Systems:
- Mac, Linux, Windows
- Programming Languages:
- Python, Shell
- Added:
- 2/9/2022
- Last Updated:
- 2/9/2022
Operations
Publications
Stringer B, de Ferrante H, Abeln S, Heringa J, Feenstra KA, Haydarlou R. PIPENN: Protein Interface Prediction with an Ensemble of Neural Nets. Unknown Journal. 2021. doi:10.1101/2021.09.03.458832.