HaForest

HaForest identifies haploinsufficient genes from epigenomic data to enable discovery of genetic contributors to diseases such as cancers and neurodevelopmental disorders.


Key Features:

  • Deep forest model: Uses a deep forest (cascade forest) architecture for tree-based ensemble learning on epigenomic features.
  • Multiscale scanning with LDA: Extracts local contextual representations from input features using a multiscale scanning approach combined with Linear Discriminant Analysis (LDA).
  • Cascade forest and feature concatenation: Integrates decision-tree-based forests in a cascade structure that concatenates features across layers to capture complex patterns and dependencies.
  • LightGBM feature extraction: Incorporates LightGBM to reveal highly expressive features by exploiting intricate dependency structures among haploinsufficient genes.
  • Robustness to bias and noise: Aims to mitigate study bias, experimental noise, and instability through its multiscale and ensemble modeling strategies.
  • Benchmarking: Validated against several computational methods and four deep learning algorithms across five epigenomic datasets.

Scientific Applications:

  • Haploinsufficiency prediction: Predicts haploinsufficient genes from epigenomic datasets.
  • Disease genetics: Supports investigation of genetic contributors to cancers and neurodevelopmental disorders linked to haploinsufficiency.
  • Epigenomic pattern discovery: Identifies epigenomic patterns and dependencies associated with haploinsufficient genes.

Methodology:

Multiscale scanning to extract local contextual representations, Linear Discriminant Analysis (LDA) for feature representation, a cascade deep forest integrating decision-tree-based forests with feature concatenation, and LightGBM for extracting expressive dependency-aware features; validated by comparisons with several computational methods and four deep learning algorithms across five epigenomic datasets.

Topics

Details

Tool Type:
command-line tool
Programming Languages:
Python
Added:
3/19/2021
Last Updated:
3/30/2021

Operations

Publications

Yang Y, Li S, Wang Y, Ma Z, Wong K, Li X. Identification of haploinsufficient genes from epigenomic data using deep forest. Briefings in Bioinformatics. 2021;22(5). doi:10.1093/bib/bbaa393. PMID:33454736.

PMID: 33454736
Funding: - National Natural Science Foundation of China: 32000464, 62076109 - Natural Science Foundation of Jilin Province: 20190103006JH - Government of the Hong Kong Special Administrative Region: 07181426 - City University of Hong Kong: CityU 11202219, CityU 11203520