SPDI
SPDI represents and normalizes genetic sequence variants using NCBI's four-attribute Sequence Position Deletion Insertion model to enable consistent normalization, projection, and aggregation of variants for genomic analyses.
Key Features:
- Four-Attribute Representation: Variants are defined by sequence, position, deletion, and insertion attributes applicable to both nucleotide and protein variants.
- Normalization and Format Conversion: Web services convert between HGVS, VCF, and SPDI formats and the NCBI Variant Overprecision Correction Algorithm generates a unique Contextual Allele for normalized representation.
- Handling Repeat Regions: The model precisely defines the reference subsequence affected by a variant, including in homopolymer and other repeat regions.
- Projection Across Congruent Sequences: Variants are projected across congruent sequences using an alignment dataset that includes non-assembly RefSeq sequences (NM, NR, NG) and inter- and intra-assembly genomic sequences (NC, NT, NW).
- Canonical Allele Designation: Among projected Contextual Alleles, one Canonical Allele is designated—typically based on the latest assembly version—to represent the set of congruent sequences for aggregation.
Scientific Applications:
- Variation databases: Representation and aggregation of asserted sequence variants in NCBI resources such as dbSNP and ClinVar.
- Genomics research: Consistent variant normalization and projection to support genomic analyses, variant curation, and the study of genetic diseases and biological functions.
Methodology:
SPDI converts variant representations between HGVS/VCF/SPDI, applies the NCBI Variant Overprecision Correction Algorithm to produce Contextual Alleles, projects variants across congruent sequences using an alignment dataset (NM/NR/NG and NC/NT/NW), and defines precise operations to address discrepancies among variant callers and databases.
Topics
Details
- Tool Type:
- api
- Added:
- 1/14/2020
- Last Updated:
- 11/24/2024
Operations
Publications
Holmes JB, Moyer E, Phan L, Maglott D, Kattman B. SPDI: data model for variants and applications at NCBI. Bioinformatics. 2019;36(6):1902-1907. doi:10.1093/bioinformatics/btz856. PMID:31738401. PMCID:PMC7523648.