t-SNE

t-SNE maps high-dimensional molecular data into a low-dimensional space to visualize and preserve local structure for characterization of molecular phenotypes.


Key Features:

  • Dimensionality Reduction: Reduces complex gene expression datasets to two or three dimensions to facilitate visualization and interpretation of sample-level patterns.
  • Data Integration: Supports integrative analysis across multiple studies and platforms for comparative analysis of molecular profiles.
  • Biological vs. Technical Variation Assessment: Enables evaluation of biological variation versus technical artifacts in sample clustering.
  • Incorporation of Additional Datasets: Permits direct incorporation of additional datasets into an existing low-dimensional representation.
  • Comparison of Molecular Subtypes: Allows comparison of molecular disease subtypes identified from separate t-SNE representations.
  • Pathway Database Integration: Enables characterization of clusters using pathway databases and supplementary data to infer underlying biological mechanisms.

Scientific Applications:

  • Elucidating Disease Mechanisms: Integrates large gene expression datasets to uncover disease mechanisms and changes in cellular pathways.
  • Patient Stratification: Stratifies patients based on molecular profiles for personalized-medicine–oriented analyses.
  • Identification of Novel Subtypes: Reveals novel molecular subtypes with distinct clinical features, including differential survival and drug responsiveness.
  • Acute Myeloid Leukemia Subtype Identification: Applied in integrative multi-omics analyses of acute myeloid leukemia to identify a myelodysplastic syndrome-like cluster and a CEBPA-mutated cluster with altered S-adenosylmethionine-dependent DNA methylation pathway activity.

Methodology:

t-SNE was applied for data-driven stratification on benchmarked multi-study, multi-platform hematological malignancy datasets, including assessment of biological versus technical variation, incorporation of additional datasets into existing low-dimensional representations, and cluster characterization using pathway databases and supplementary data.

Topics

Collections

Details

Tool Type:
command-line tool
Programming Languages:
C++, R
Added:
1/20/2021
Last Updated:
5/21/2021

Operations

Publications

Mehtonen J, Pölönen P, Häyrynen S, Dufva O, Lin J, Liuksiala T, Granberg K, Lohi O, Hautamäki V, Nykter M, Heinäniemi M. Data-driven characterization of molecular phenotypes across heterogeneous sample collections. Nucleic Acids Research. 2019;47(13):e76-e76. doi:10.1093/nar/gkz281. PMID:31329928. PMCID:PMC6648337.

PMID: 31329928
PMCID: PMC6648337
Funding: - Academy of Finland: 269474, 276634