ExpressionBlast

ExpressionBlast performs large-scale similarity search and comparative analysis of gene expression profiles across heterogeneous public repositories to identify related perturbations, pathways, phenotypes, and disease states.


Key Features:

  • Metadata extraction and classification: Automates metadata extraction, replicate identification, and treatment–control classification using text mining and heuristic matching.
  • Probe harmonization: Harmonizes probe identifiers across experiments and platforms.
  • Expression scaling and normalization: Converts expression values to consistent log-ratio scales and normalizes distributions.
  • Large GEO index: Indexes >900,000 samples from >40,000 Gene Expression Omnibus (GEO) series.
  • Similarity scoring: Computes similarity scores using distance metrics such as Euclidean distance, correlation, and anti-correlation.
  • Cross-species orthology: Incorporates orthology mappings via InParanoid and resolves ambiguous mappings by closest expression similarity.
  • Statistical augmentation and enrichment: Augments matches with statistical significance estimates, GO term enrichment, and key shared text-mined keywords, and links to supporting publication abstracts.
  • Visualization: Visualizes results as heatmaps.

Scientific Applications:

  • Perturbation and phenotype discovery: Identifies related perturbations, pathways, phenotypes, and disease states by matching expression signatures across datasets.
  • Cross-platform and cross-species comparison: Enables comparative analysis across platforms and species using probe harmonization and InParanoid orthology mappings.
  • Functional genomics and systems biology: Supports hypothesis generation and interpretation in functional genomics and systems biology by mining public expression archives.
  • Meta-analysis and literature linking: Facilitates meta-analysis of GEO datasets and links expression signatures to supporting publication abstracts and enriched GO terms.

Methodology:

Metadata extraction, replicate identification, and treatment–control classification via text mining and heuristic matching; harmonization of probe identifiers; conversion of expression values to log-ratio scales and distribution normalization; indexing of >900,000 GEO samples; similarity scoring using Euclidean distance, correlation, and anti-correlation; orthology mapping via InParanoid with ambiguous mappings resolved by closest expression similarity; computation of statistical significance estimates, GO enrichment, and extraction of shared text-mined keywords with links to publication abstracts.

Topics

Details

Tool Type:
command-line tool
Added:
12/18/2017
Last Updated:
11/25/2024

Operations

Publications

Zinman GE, Naiman S, Kanfi Y, Cohen H, Bar-Joseph Z. ExpressionBlast: mining large, unstructured expression databases. Nature Methods. 2013;10(10):925-926. doi:10.1038/nmeth.2630. PMID:24076985.

Links