iDHS-DASTS

iDHS-DASTS identifies DNase I hypersensitive sites (DHS) from DNA sequences to detect regulatory regions involved in gene regulation.


Key Features:

  • Feature Extraction with PseDNC: Employs Pseudo Dinucleotide Composition (PseDNC) to capture intrinsic DNA properties and spatial information from sequences.
  • Dimensionality Reduction via LASSO: Uses the Least Absolute Shrinkage and Selection Operator (LASSO) to reduce feature dimensionality while selecting informative features.
  • Stacking Learning Classifier: Implements a stacking ensemble integrating Adaboost, Random Forest, Gradient Boosting, Extra Trees, and Support Vector Machine (SVM) to combine multiple classifiers for prediction.
  • Handling Imbalanced Datasets (SMOTE-Tomek): Applies SMOTE-Tomek to balance class distributions prior to training.

Scientific Applications:

  • Benchmark DHS identification: Demonstrated accuracies of 92.06% (S1), 91.06% (S2), 90.72% (S3), and 90.31% on independent dataset S4.

Methodology:

Feature extraction using PseDNC, dimensionality reduction with LASSO, data balancing via SMOTE-Tomek, and classification by a stacking ensemble combining Adaboost, Random Forest, Gradient Boosting, Extra Trees, and SVM.

Topics

Details

Tool Type:
api
Added:
1/18/2021
Last Updated:
2/2/2021

Operations

Publications

Zhang S, Duan Z, Yang W, Qian C, You Y. iDHS-DASTS: identifying DNase I hypersensitive sites based on LASSO and stacking learning. Molecular Omics. 2021;17(1):130-141. doi:10.1039/d0mo00115e. PMID:33295914.

PMID: 33295914
Funding: - National Natural Science Foundation of China: 11601407