iDHS-DASTS
iDHS-DASTS identifies DNase I hypersensitive sites (DHS) from DNA sequences to detect regulatory regions involved in gene regulation.
Key Features:
- Feature Extraction with PseDNC: Employs Pseudo Dinucleotide Composition (PseDNC) to capture intrinsic DNA properties and spatial information from sequences.
- Dimensionality Reduction via LASSO: Uses the Least Absolute Shrinkage and Selection Operator (LASSO) to reduce feature dimensionality while selecting informative features.
- Stacking Learning Classifier: Implements a stacking ensemble integrating Adaboost, Random Forest, Gradient Boosting, Extra Trees, and Support Vector Machine (SVM) to combine multiple classifiers for prediction.
- Handling Imbalanced Datasets (SMOTE-Tomek): Applies SMOTE-Tomek to balance class distributions prior to training.
Scientific Applications:
- Benchmark DHS identification: Demonstrated accuracies of 92.06% (S1), 91.06% (S2), 90.72% (S3), and 90.31% on independent dataset S4.
Methodology:
Feature extraction using PseDNC, dimensionality reduction with LASSO, data balancing via SMOTE-Tomek, and classification by a stacking ensemble combining Adaboost, Random Forest, Gradient Boosting, Extra Trees, and SVM.
Topics
Details
- Tool Type:
- api
- Added:
- 1/18/2021
- Last Updated:
- 2/2/2021
Operations
Publications
Zhang S, Duan Z, Yang W, Qian C, You Y. iDHS-DASTS: identifying DNase I hypersensitive sites based on LASSO and stacking learning. Molecular Omics. 2021;17(1):130-141. doi:10.1039/d0mo00115e. PMID:33295914.
DOI: 10.1039/D0MO00115E
PMID: 33295914
Funding: - National Natural Science Foundation of China: 11601407