SPDI

SPDI represents and normalizes genetic sequence variants using NCBI's four-attribute Sequence Position Deletion Insertion model to enable consistent normalization, projection, and aggregation of variants for genomic analyses.


Key Features:

  • Four-Attribute Representation: Variants are defined by sequence, position, deletion, and insertion attributes applicable to both nucleotide and protein variants.
  • Normalization and Format Conversion: Web services convert between HGVS, VCF, and SPDI formats and the NCBI Variant Overprecision Correction Algorithm generates a unique Contextual Allele for normalized representation.
  • Handling Repeat Regions: The model precisely defines the reference subsequence affected by a variant, including in homopolymer and other repeat regions.
  • Projection Across Congruent Sequences: Variants are projected across congruent sequences using an alignment dataset that includes non-assembly RefSeq sequences (NM, NR, NG) and inter- and intra-assembly genomic sequences (NC, NT, NW).
  • Canonical Allele Designation: Among projected Contextual Alleles, one Canonical Allele is designated—typically based on the latest assembly version—to represent the set of congruent sequences for aggregation.

Scientific Applications:

  • Variation databases: Representation and aggregation of asserted sequence variants in NCBI resources such as dbSNP and ClinVar.
  • Genomics research: Consistent variant normalization and projection to support genomic analyses, variant curation, and the study of genetic diseases and biological functions.

Methodology:

SPDI converts variant representations between HGVS/VCF/SPDI, applies the NCBI Variant Overprecision Correction Algorithm to produce Contextual Alleles, projects variants across congruent sequences using an alignment dataset (NM/NR/NG and NC/NT/NW), and defines precise operations to address discrepancies among variant callers and databases.

Topics

Details

Tool Type:
api
Added:
1/14/2020
Last Updated:
11/24/2024

Operations

Publications

Holmes JB, Moyer E, Phan L, Maglott D, Kattman B. SPDI: data model for variants and applications at NCBI. Bioinformatics. 2019;36(6):1902-1907. doi:10.1093/bioinformatics/btz856. PMID:31738401. PMCID:PMC7523648.