iRefIndex
iRefIndex integrates protein–protein interaction data from multiple primary databases into a non-redundant index using sequence- and taxonomy-based keys to support comparative and systems-level analysis of interaction networks.
Key Features:
- Unification of Databases: Consolidates interaction data from BIND, BioGRID, DIP, HPRD, IntAct, MINT, MPact, MPPI, and OPHID into a single indexed resource.
- Redundancy Resolution: Generates unique keys for interaction records and participant proteins using primary sequences and taxonomy identifiers combined with the Secure Hash Algorithm (SHA) to group identical records across sources.
- Mapping and Feedback Mechanism: Maps protein database references to current protein sequence records and assigns a mapping score that identifies malformed, deprecated, ambiguous, or unfound references and provides feedback to source databases.
- Independent Retrieval: Produces participant-protein keys that enable retrieval of interaction information independently of the original protein reference identifiers used in source records.
- Flexible Grouping Options: Supports alternative redundant groupings based on gene identifiers or near-sequence identity in addition to exact sequence-based grouping.
- Standard Output Format: Provides the indexed interaction data in PSI-MITAB 2.5 format.
Scientific Applications:
- Interaction dataset curation: Resolves redundant and inconsistent records to improve the quality and consistency of aggregated interaction datasets.
- Comparative interactomics: Enables cross-database comparisons and identification of conserved or database-specific interactions across species using sequence- and taxonomy-based grouping.
- Systems biology: Facilitates construction and analysis of protein interaction networks for systems-level investigations.
- Drug discovery: Supports compilation and evaluation of protein interaction partners relevant to target identification and validation.
- Functional genomics: Aids inference of protein function and functional modules by providing consolidated interaction evidence.
Methodology:
Unique keys for interaction records and participant proteins are generated from primary sequences and taxonomy identifiers using the Secure Hash Algorithm (SHA); protein database references are mapped to current sequence records and assigned a mapping score to flag malformed, deprecated, ambiguous, or unfound references, and identical records are grouped based on sequence and taxonomy information.
Topics
Details
- Maturity:
- Mature
- Tool Type:
- web application
- Operating Systems:
- Linux, Windows, Mac
- Added:
- 3/4/2016
- Last Updated:
- 11/25/2024
Operations
Data Inputs & Outputs
Query and retrieval
Inputs
Outputs
Publications
Razick S, Magklaras G, Donaldson IM. iRefIndex: A consolidated protein interaction database with provenance. BMC Bioinformatics. 2008;9(1). doi:10.1186/1471-2105-9-405. PMID:18823568. PMCID:PMC2573892.