Harvestman

Harvestman performs hierarchical feature learning and selection from whole genome sequencing (WGS) data to generate biologically informed feature representations for genomic classification and interpretation.


Key Features:

  • Hierarchical Feature Learning: Exploits hierarchical relationships among genomic variant representations to perform automatic feature learning that captures complex biological structure.
  • Scalability: Processes large-scale WGS datasets, demonstrated on 1000 Genomes Project phase 3 data containing over 84 million variants and thousands of genomes.
  • Adaptive Feature Selection: Selects and combines diverse feature representations tailored to specific learning tasks, as shown in analyses of TCGA breast cancer data.
  • Efficient Feature Subset Selection: Identifies smaller, less redundant feature subsets while maintaining or improving classifier accuracy relative to traditional binary SNP representation methods.

Scientific Applications:

  • Genomic Variant Analysis: Enables population-scale variant analysis using datasets such as the 1000 Genomes Project phase 3 to study human genetic diversity and its implications for health.
  • Cancer Research: Applied to The Cancer Genome Atlas (TCGA) breast cancer data to identify genomic features associated with cancer phenotypes.

Methodology:

Harvestman integrates hierarchical relationships among genomic variants into the feature learning and selection process, producing nuanced representations that facilitate improved model performance and interpretability.

Topics

Details

Tool Type:
command-line tool
Added:
1/18/2021
Last Updated:
1/30/2021

Operations

Publications

Frisby TS, Baker SJ, Marçais G, Hoang QM, Kingsford C, Langmead CJ. Harvestman: A framework for hierarchical feature learning and selection from whole genome sequencing data. Unknown Journal. 2020. doi:10.1101/2020.03.24.005603.