H-InvDB

H-InvDB provides comprehensive annotations of human genes and transcripts for genetic and functional genomics research (version 6.2: 219,765 transcripts organized into 43,159 gene clusters) derived from full-length cDNAs and mRNAs.


Key Features:

  • Gene structures: Detailed annotation of exon–intron organization and gene models.
  • Alternative splicing isoforms: Annotation and cataloguing of transcript isoforms produced by alternative splicing.
  • Non-coding RNAs (ncRNAs): Annotation of non-coding functional RNAs and relationships to known ncRNAs.
  • Protein functions: Functional annotation of encoded proteins based on sequence and database evidence.
  • Functional domains: Identification and annotation of protein functional domains.
  • Sub-cellular localization: Annotation of predicted or known protein sub-cellular localizations.
  • Metabolic pathways: Assignment of proteins to metabolic pathways.
  • Three-dimensional protein structures: Links or annotations related to 3D protein structure information.
  • Genetic polymorphisms: Annotation of single nucleotide polymorphisms (SNPs), insertions/deletions (indels), and microsatellite repeats.
  • Disease associations: Annotation of gene and transcript associations with human diseases.
  • Gene expression profiles: Inclusion of expression profile annotations for transcripts and genes.
  • Molecular evolutionary features: Annotation of evolutionary characteristics of genes and proteins.
  • Protein–protein interactions: Annotation of known or inferred protein–protein interactions.
  • Gene families/groups: Classification of genes into families and groups.
  • Microarray probe mapping: Mapping of microarray probes to transcripts and gene models.
  • Updated gene models: Inclusion of revised and updated gene model annotations.
  • Glycogene integration: Inclusion of data from the Glycogene database.
  • Navigation search: Query capability combining up to 16 options for complex retrieval criteria.
  • H-InvDB Enrichment Analysis Tool (HEAT): Detection of annotations significantly enriched in user-defined gene sets compared with the entire database of representative transcripts.
  • Web service APIs: Programmatic access via SOAP and REST web service interfaces.

Scientific Applications:

  • Gene model curation: Refinement and validation of exon–intron structures and transcript models.
  • Alternative splicing analysis: Study of transcript isoform diversity and splicing patterns.
  • ncRNA characterization: Identification and characterization of non-coding functional RNAs and their relationships to genes.
  • Protein function and domain analysis: Functional annotation and domain-based analyses of protein products.
  • Sub-cellular localization studies: Investigation of protein localization in cellular compartments.
  • Pathway reconstruction: Assignment of genes and proteins to metabolic and signaling pathways.
  • Structural biology integration: Linking transcript/ protein annotations to three-dimensional structure data.
  • Genetic variation and disease association studies: Analysis of SNPs, indels, microsatellites and their associations with disease.
  • Expression profiling: Interpretation and re-annotation of gene expression datasets, including microarray probe mapping.
  • Evolutionary and comparative genomics: Analyses of molecular evolutionary features across gene families.
  • Protein–protein interaction analysis: Exploration of interaction networks informed by annotated interactions.
  • Gene set enrichment analysis: Identification of annotations enriched in user-defined gene lists using HEAT.
  • Programmatic integration: Integration of H-InvDB data into computational workflows via SOAP and REST APIs.

Methodology:

Annotations are derived from full-length cDNAs and mRNAs; gene models are updated and microarray probes are mapped; relationships to known ncRNAs and Glycogene database data are integrated; HEAT detects annotations significantly enriched in user-defined gene sets relative to the database of representative transcripts; web service APIs are implemented using SOAP and REST.

Topics

Collections

Details

Tool Type:
web application
Operating Systems:
Linux, Windows, Mac
Added:
9/19/2015
Last Updated:
11/24/2024

Operations

Data Inputs & Outputs

Other operations do not define inputs or outputs.

Publications

Yamasaki C, Murakami K, Takeda J, Sato Y, Noda A, Sakate R, Habara T, Nakaoka H, Todokoro F, Matsuya A, Imanishi T, Gojobori T. H-InvDB in 2009: extended database and data mining resources for human genes and transcripts. Nucleic Acids Research. 2009;38(suppl_1):D626-D632. doi:10.1093/nar/gkp1020. PMID:19933760. PMCID:PMC2808976.

Documentation