Hi-LASSO
Hi-LASSO applies a refined LASSO algorithm implemented for Python and Apache Spark to perform feature selection in high-dimensional biological datasets using parallel computing and parametric statistical tests to improve selection accuracy.
Key Features:
- High-Performance Parallel Computing: Implements parallel processing in Python and Apache Spark to reduce computation time for high-dimensional datasets.
- Superior Feature Selection: Demonstrates improved feature selection performance over other LASSO variants, aiding identification of relevant genes or biomarkers.
- Parametric Statistical Tests: Incorporates parametric statistical tests to validate selected features and provide a rigorous statistical framework for selection.
- Large-Scale Data Processing: Integrates with Apache Spark to support processing of large-scale bioinformatics datasets.
Scientific Applications:
- Genomics: Selection of relevant genes and genetic markers from high-dimensional genomic datasets.
- Proteomics: Identification of proteomic features and candidate protein biomarkers in high-dimensional proteomics data.
- Omics biomarker discovery: Extraction of meaningful patterns and biomarkers associated with diseases, genetic traits, or environmental interactions across omics studies.
Methodology:
Uses a refined LASSO (Least Absolute Shrinkage and Selection Operator) algorithm optimized for high-dimensional datasets combined with parallel processing and parametric statistical tests, provided as a Python package and an Apache Spark library.
Topics
Details
- License:
- MIT
- Cost:
- Free of charge
- Tool Type:
- library
- Operating Systems:
- Mac, Linux, Windows
- Programming Languages:
- Python
- Added:
- 1/22/2023
- Last Updated:
- 11/24/2024
Operations
Publications
Jo J, Jung S, Park J, Kim Y, Kang M. Hi-LASSO: High-performance python and apache spark packages for feature selection with high-dimensional data. PLOS ONE. 2022;17(12):e0278570. doi:10.1371/journal.pone.0278570. PMID:36455001. PMCID:PMC9714948.