Hi-LASSO

Hi-LASSO applies a refined LASSO algorithm implemented for Python and Apache Spark to perform feature selection in high-dimensional biological datasets using parallel computing and parametric statistical tests to improve selection accuracy.


Key Features:

  • High-Performance Parallel Computing: Implements parallel processing in Python and Apache Spark to reduce computation time for high-dimensional datasets.
  • Superior Feature Selection: Demonstrates improved feature selection performance over other LASSO variants, aiding identification of relevant genes or biomarkers.
  • Parametric Statistical Tests: Incorporates parametric statistical tests to validate selected features and provide a rigorous statistical framework for selection.
  • Large-Scale Data Processing: Integrates with Apache Spark to support processing of large-scale bioinformatics datasets.

Scientific Applications:

  • Genomics: Selection of relevant genes and genetic markers from high-dimensional genomic datasets.
  • Proteomics: Identification of proteomic features and candidate protein biomarkers in high-dimensional proteomics data.
  • Omics biomarker discovery: Extraction of meaningful patterns and biomarkers associated with diseases, genetic traits, or environmental interactions across omics studies.

Methodology:

Uses a refined LASSO (Least Absolute Shrinkage and Selection Operator) algorithm optimized for high-dimensional datasets combined with parallel processing and parametric statistical tests, provided as a Python package and an Apache Spark library.

Topics

Details

License:
MIT
Cost:
Free of charge
Tool Type:
library
Operating Systems:
Mac, Linux, Windows
Programming Languages:
Python
Added:
1/22/2023
Last Updated:
11/24/2024

Operations

Publications

Jo J, Jung S, Park J, Kim Y, Kang M. Hi-LASSO: High-performance python and apache spark packages for feature selection with high-dimensional data. PLOS ONE. 2022;17(12):e0278570. doi:10.1371/journal.pone.0278570. PMID:36455001. PMCID:PMC9714948.

PMID: 36455001
PMCID: PMC9714948
Funding: - National Research Foundation of Korea: NRF-2021R1I1A3048029

Documentation

Links