EMU

EMU infers population structure from large-scale genetic datasets with substantial missing genotype data.


Key Features:

  • Handling Missing Data: Manages non-random missingness in genotype data to prevent bias in principal component analyses.
  • EM-PCA algorithm: Implements an Expectation-Maximization Principal Component Analysis (EM-PCA) approach to infer population structure despite rampant missingness.
  • Performance and scalability: Shows superior accuracy and competitive computational efficiency relative to alternative PCA methods based on extensive simulations.
  • Large-cohort capability: Recovered Han Chinese population structure in the Chinese Millionome Project Phase 1 dataset (~100,000 individuals at ~0.08× coverage) using CPU hours.

Scientific Applications:

  • Ultra-low coverage sequencing: Infer population structure from datasets generated by non-invasive prenatal tests and other ultra-low coverage sequencing protocols.
  • Shallowly sequenced cohorts: Analyze population structure in large shallow-sequenced projects such as the Chinese Millionome Project Phase 1.
  • Population-genetics analyses: Provide population structure estimates required for downstream genetic and genomic analyses that must account for missing genotype data.

Methodology:

Uses an Expectation-Maximization Principal Component Analysis (EM-PCA) algorithm and evaluation via extensive simulations comparing to other PCA methods.

Topics

Details

License:
GPL-3.0
Tool Type:
command-line tool
Programming Languages:
Python
Added:
3/19/2021
Last Updated:
5/5/2021

Operations

Publications

Meisner J, Liu S, Huang M, Albrechtsen A. Large-scale inference of population structure in presence of missingness using PCA. Bioinformatics. 2021;37(13):1868-1875. doi:10.1093/bioinformatics/btab027. PMID:33459779.

PMID: 33459779
Funding: - Lundbeck foundation: R215-2015-4174 - National Natural Science Foundation of China: 31900487