Harvestman
Harvestman performs hierarchical feature learning and selection from whole genome sequencing (WGS) data to generate biologically informed feature representations for genomic classification and interpretation.
Key Features:
- Hierarchical Feature Learning: Exploits hierarchical relationships among genomic variant representations to perform automatic feature learning that captures complex biological structure.
- Scalability: Processes large-scale WGS datasets, demonstrated on 1000 Genomes Project phase 3 data containing over 84 million variants and thousands of genomes.
- Adaptive Feature Selection: Selects and combines diverse feature representations tailored to specific learning tasks, as shown in analyses of TCGA breast cancer data.
- Efficient Feature Subset Selection: Identifies smaller, less redundant feature subsets while maintaining or improving classifier accuracy relative to traditional binary SNP representation methods.
Scientific Applications:
- Genomic Variant Analysis: Enables population-scale variant analysis using datasets such as the 1000 Genomes Project phase 3 to study human genetic diversity and its implications for health.
- Cancer Research: Applied to The Cancer Genome Atlas (TCGA) breast cancer data to identify genomic features associated with cancer phenotypes.
Methodology:
Harvestman integrates hierarchical relationships among genomic variants into the feature learning and selection process, producing nuanced representations that facilitate improved model performance and interpretability.
Topics
Details
- Tool Type:
- command-line tool
- Added:
- 1/18/2021
- Last Updated:
- 1/30/2021
Operations
Publications
Frisby TS, Baker SJ, Marçais G, Hoang QM, Kingsford C, Langmead CJ. Harvestman: A framework for hierarchical feature learning and selection from whole genome sequencing data. Unknown Journal. 2020. doi:10.1101/2020.03.24.005603.