ExpressionBlast
ExpressionBlast performs large-scale similarity search and comparative analysis of gene expression profiles across heterogeneous public repositories to identify related perturbations, pathways, phenotypes, and disease states.
Key Features:
- Metadata extraction and classification: Automates metadata extraction, replicate identification, and treatment–control classification using text mining and heuristic matching.
- Probe harmonization: Harmonizes probe identifiers across experiments and platforms.
- Expression scaling and normalization: Converts expression values to consistent log-ratio scales and normalizes distributions.
- Large GEO index: Indexes >900,000 samples from >40,000 Gene Expression Omnibus (GEO) series.
- Similarity scoring: Computes similarity scores using distance metrics such as Euclidean distance, correlation, and anti-correlation.
- Cross-species orthology: Incorporates orthology mappings via InParanoid and resolves ambiguous mappings by closest expression similarity.
- Statistical augmentation and enrichment: Augments matches with statistical significance estimates, GO term enrichment, and key shared text-mined keywords, and links to supporting publication abstracts.
- Visualization: Visualizes results as heatmaps.
Scientific Applications:
- Perturbation and phenotype discovery: Identifies related perturbations, pathways, phenotypes, and disease states by matching expression signatures across datasets.
- Cross-platform and cross-species comparison: Enables comparative analysis across platforms and species using probe harmonization and InParanoid orthology mappings.
- Functional genomics and systems biology: Supports hypothesis generation and interpretation in functional genomics and systems biology by mining public expression archives.
- Meta-analysis and literature linking: Facilitates meta-analysis of GEO datasets and links expression signatures to supporting publication abstracts and enriched GO terms.
Methodology:
Metadata extraction, replicate identification, and treatment–control classification via text mining and heuristic matching; harmonization of probe identifiers; conversion of expression values to log-ratio scales and distribution normalization; indexing of >900,000 GEO samples; similarity scoring using Euclidean distance, correlation, and anti-correlation; orthology mapping via InParanoid with ambiguous mappings resolved by closest expression similarity; computation of statistical significance estimates, GO enrichment, and extraction of shared text-mined keywords with links to publication abstracts.
Topics
Details
- Tool Type:
- command-line tool
- Added:
- 12/18/2017
- Last Updated:
- 11/25/2024
Operations
Publications
Zinman GE, Naiman S, Kanfi Y, Cohen H, Bar-Joseph Z. ExpressionBlast: mining large, unstructured expression databases. Nature Methods. 2013;10(10):925-926. doi:10.1038/nmeth.2630. PMID:24076985.