ropls
ropls performs multivariate statistical analysis of omics datasets, implementing Principal Component Analysis (PCA), Partial Least Squares (PLS), and Orthogonal Partial Least Squares (OPLS) to model predictive and orthogonal variation in high-dimensional, multicollinear data.
Key Features:
- Principal Component Analysis (PCA): Unsupervised dimensionality reduction for visualization and exploratory analysis of multivariate omics data.
- Partial Least Squares (PLS): Supervised regression and classification modeling linking predictor variables to outcomes.
- Orthogonal Partial Least Squares (OPLS): Decomposition of variation into predictive (correlated with the factor of interest) and orthogonal (uncorrelated) components.
- High-dimensional data handling: Designed for datasets with more variables than samples and substantial multicollinearity among variables.
- Variable Importance in Projection (VIP): Multivariate scoring of feature relevance for PLS/OPLS models.
- Univariate hypothesis testing: Integration of univariate p-values alongside multivariate scores for feature assessment.
- Correlation analysis and clustering: Tools for correlation-based cohort stratification and cluster identification.
- Visualization: Generation of score and loading representations to interpret multivariate models.
- Feature selection: Combined univariate and multivariate criteria to prioritize metabolites or other variables.
- Application to LC-HRMS metabolomics: Demonstrated use with liquid chromatography–high-resolution mass spectrometry (LC-HRMS) urine metabolomics datasets.
Scientific Applications:
- Exploratory data analysis: Visualizing structure and major sources of variation in omics datasets using PCA and OPLS score plots.
- Regression and classification: Building PLS/OPLS models to predict continuous outcomes or classify sample groups.
- Feature selection and biomarker discovery: Identifying metabolites of interest by combining VIP scores and univariate p-values.
- Metabolomics cohort analysis: Assessing effects of physiological factors such as age, body mass index (BMI), and gender on the urinary metabolome.
- Cohort stratification: Using correlation analysis and clustering to define subgroups within a cohort (e.g., stratifying 183 adults and identifying 108 differential urine metabolites).
Methodology:
PCA, PLS, OPLS, univariate hypothesis testing, variable importance in projection (VIP) scoring, correlation analysis and clustering, and integration of univariate p-values with multivariate VIP.
Topics
Collections
Details
- License:
- CECILL-2.1
- Tool Type:
- command-line tool, library
- Operating Systems:
- Linux, Windows, Mac
- Programming Languages:
- R
- Added:
- 1/17/2017
- Last Updated:
- 1/11/2019
Operations
Publications
Thévenot EA, Roux A, Xu Y, Ezan E, Junot C. Analysis of the Human Adult Urinary Metabolome Variations with Age, Body Mass Index, and Gender by Implementing a Comprehensive Workflow for Univariate and OPLS Statistical Analyses. Journal of Proteome Research. 2015;14(8):3322-3335. doi:10.1021/acs.jproteome.5b00354. PMID:26088811.