NCBI reference sequences
NCBI reference sequences provide a curated, non-redundant collection of genomic, transcript, and protein reference records selected from public sequence archives with structured annotations for genome annotation, comparative genomics, and variant analysis.
Key Features:
- Non-redundant reference set: Records are selected from public sequence archives to produce an integrated, non-redundant collection of genomic, transcript, and protein sequences.
- Extensive taxonomic coverage: The collection covers sequences from over 55,000 organisms, including more than 4,800 viruses, ~40,000 prokaryotes, and ~10,000 eukaryotes.
- Curated annotations: Annotations include coding regions, conserved domains, tRNAs, sequence-tagged sites (STS), variations, literature references, gene and protein product names, and cross-references to other databases.
- Annotation methods: Annotations are produced through automated analyses, prediction algorithms, propagation from GenBank, and manual curation by NCBI staff.
- Quality assurance: Records undergo format checks and quality tests for sequence and annotation accuracy.
- Prokaryotic Genome Annotation Pipeline (PGAP): PGAP combines alignment-based methods with ab initio gene predictions for prokaryotic genome annotation.
- Viral genomics resource: The NCBI Viral Genomes Resource catalogs and curates reference viral genome sequences and leverages community knowledge for well-annotated viral reference sets.
- Taxonomic validation and expansion: Ongoing taxonomic validation and expansion support genome annotation, comparative genomics, and clinical testing applications.
Scientific Applications:
- Genome annotation: Provides stable reference sequences and annotations for annotating genomes across diverse organisms.
- Gene identification and characterization: Supplies curated sequences and feature annotations to support discovery and functional characterization of genes.
- Mutation and polymorphism analysis: Enables identification and interpretation of genetic variations against reference sequences.
- Expression studies: Supports analysis of transcript-level data and gene expression by providing transcript reference sequences and annotations.
- Comparative genomics: Facilitates comparative analyses to assess evolutionary relationships and functional conservation across taxa.
- Viral genome reference curation: Provides curated viral reference genomes to support viral genomics and related analyses.
Methodology:
Records are selected from public sequence archives and processed using automated analyses, prediction algorithms, propagation from GenBank, manual curation by NCBI staff, format checks and quality tests, alignment-based methods and ab initio gene predictions (as implemented in PGAP), and taxonomic validation.
Topics
Collections
Details
- Tool Type:
- web application
- Operating Systems:
- Linux, Windows, Mac
- Added:
- 10/7/2015
- Last Updated:
- 11/24/2024
Operations
Publications
Pruitt KD, Tatusova T, Brown GR, Maglott DR. NCBI Reference Sequences (RefSeq): current status, new features and genome annotation policy. Nucleic Acids Research. 2011;40(D1):D130-D135. doi:10.1093/nar/gkr1079. PMID:22121212. PMCID:PMC3245008.
Pruitt KD, Tatusova T, Maglott DR. NCBI reference sequences (RefSeq): a curated non-redundant sequence database of genomes, transcripts and proteins. Nucleic Acids Research. 2007;35(Database):D61-D65. doi:10.1093/nar/gkl842. PMID:17130148. PMCID:PMC1716718.
O'Leary NA, Wright MW, Brister JR, Ciufo S, Haddad D, McVeigh R, Rajput B, Robbertse B, Smith-White B, Ako-Adjei D, Astashyn A, Badretdin A, Bao Y, Blinkova O, Brover V, Chetvernin V, Choi J, Cox E, Ermolaeva O, Farrell CM, Goldfarb T, Gupta T, Haft D, Hatcher E, Hlavina W, Joardar VS, Kodali VK, Li W, Maglott D, Masterson P, McGarvey KM, Murphy MR, O'Neill K, Pujar S, Rangwala SH, Rausch D, Riddick LD, Schoch C, Shkeda A, Storz SS, Sun H, Thibaud-Nissen F, Tolstoy I, Tully RE, Vatsan AR, Wallin C, Webb D, Wu W, Landrum MJ, Kimchi A, Tatusova T, DiCuccio M, Kitts P, Murphy TD, Pruitt KD. Reference sequence (RefSeq) database at NCBI: current status, taxonomic expansion, and functional annotation. Nucleic Acids Research. 2015;44(D1):D733-D745. doi:10.1093/nar/gkv1189. PMID:26553804. PMCID:PMC4702849.
Tatusova T, DiCuccio M, Badretdin A, Chetvernin V, Nawrocki EP, Zaslavsky L, Lomsadze A, Pruitt KD, Borodovsky M, Ostell J. NCBI prokaryotic genome annotation pipeline. Nucleic Acids Research. 2016;44(14):6614-6624. doi:10.1093/nar/gkw569. PMID:27342282. PMCID:PMC5001611.
Brister JR, Ako-adjei D, Bao Y, Blinkova O. NCBI Viral Genomes Resource. Nucleic Acids Research. 2014;43(D1):D571-D577. doi:10.1093/nar/gku1207. PMID:25428358. PMCID:PMC4383986.