GAEseq
GAEseq reconstructs genomic mixtures from high-throughput DNA sequencing data for single individual haplotyping and viral quasispecies reconstruction using a graph auto-encoder.
Key Features:
- Graph auto-encoder framework: Leverages structural properties inherent in sequencing data to model relationships between reads and variants.
- Neural network-based inference: Trains a neural network to mitigate sequencing errors and infer posterior probabilities of sequencing read origins.
- Read-origin assignment: Formulates and addresses the NP-hard problem of assigning reads to mixture components.
- Consensus-based reconstruction: Identifies and reconstructs mixture components by achieving consensus among reads inferred to originate from the same genomic component.
- Benchmark performance: Evaluated on realistic synthetic data and experimental datasets and reported to outperform state-of-the-art methods in haplotype assembly and viral community reconstruction.
Scientific Applications:
- Single individual haplotyping: Assembly of haplotypes from sequencing reads to resolve diploid or polyploid genomes.
- Viral quasispecies reconstruction: Reconstruction of viral community composition and strain sequences from mixed viral populations.
- Genomic mixture reconstruction: Deconvolution of mixed genomic samples into constituent components using read-origin inference.
Methodology:
Applies a graph auto-encoder neural network trained to infer posterior probabilities of read origins while mitigating sequencing errors and reconstructs mixture components by consensus among reads assigned to the same component.
Topics
Details
- Tool Type:
- command-line tool
- Programming Languages:
- C++, Python
- Added:
- 1/14/2020
- Last Updated:
- 1/14/2021
Operations
Publications
Ke Z, Vikalo H. A Graph Auto-Encoder for Haplotype Assembly and Viral Quasispecies Reconstruction. Unknown Journal. 2019. doi:10.1101/837674.
Ke Z, Vikalo H. A Graph Auto-Encoder for Haplotype Assembly and Viral Quasispecies Reconstruction. Proceedings of the AAAI Conference on Artificial Intelligence. 2020;34(01):719-726. doi:10.1609/aaai.v34i01.5414.