PHG

PHG constructs and queries a pangenome haplotype graph to store haplotypes and variant information and to impute missing genotypes for genome-wide imputation and genomic prediction in breeding programs.


Key Features:

  • Pangenome database: Stores all identified haplotypes and variant information across a set of individuals.
  • PHG construction at multiple scales: Supports PHG builds comprising different sample sets, exemplified by a 24-individual and a 398-individual PHG.
  • Genic-region diversity capture: Captures genetic diversity present within genic regions of the sorghum genome.
  • Low-coverage SNP calling: Calls single nucleotide polymorphisms (SNPs) accurately at very low sequencing coverage (reported error rate 5.9% at 0.01x) with performance near that at higher coverage (8x).
  • Genotype imputation: Imputes missing genotypes in progeny using parental haplotypes stored within the PHG.
  • Cross-platform genotype unification: Unifies genotype calls from different sequencing platforms and methods, including genotyping-by-sequencing (GBS) and rhAmpSeq.
  • Efficient storage: Stores sequence and variant data in an efficient format to reduce input sequence requirements and genotyping costs.
  • Genomic prediction support: Enables genomic prediction using imputed genotypes with reported mean accuracies between 0.57 and 0.73 for various traits.

Scientific Applications:

  • Genomic selection in breeding programs: Applied to the Chibas sorghum breeding program by sequencing 24 founders at low coverage and imputing genotypes in progeny for selection.
  • Genotype imputation in progeny: Imputes genotypes in 207 progeny using parental haplotypes stored in the PHG.
  • Comparative genotyping: Provides genotype datasets comparable to genotyping-by-sequencing (GBS) and sequencing targeted amplicons (rhAmpSeq) for downstream analyses.
  • Cost-effective genotyping for large populations: Reduces required sequencing input to lower genotyping costs and enable larger breeding populations to capture genetic recombination.
  • Cross-species pangenome storage: Maintains comprehensive variant information applicable to diverse taxa and any species.

Methodology:

Construct PHG pangenome databases storing haplotypes and variants (examples: 24- and 398-individual PHGs), perform SNP calling from low-coverage sequencing (reported 0.01x and 8x comparisons), and impute missing progeny genotypes using parental haplotypes stored in the PHG, as demonstrated by processing 24 low-coverage founders and imputing 207 progeny.

Topics

Details

Added:
1/9/2020
Last Updated:
1/9/2021

Operations

Publications

Jensen SE, Charles JR, Muleta K, Bradbury P, Casstevens T, Deshpande SP, Gore MA, Gupta R, Ilut DC, Johnson L, Lozano R, Miller Z, Ramu P, Rathore A, Romay MC, Upadhyaya HD, Varshney R, Morris GP, Pressoir G, Buckler ES, Ramstein GP. A sorghum Practical Haplotype Graph facilitates genome-wide imputation and cost-effective genomic prediction. Unknown Journal. 2019. doi:10.1101/775221.