cNMTF

cNMTF integrates genotype, phenotype, variant-effect, and gene-network data using corrected non-negative matrix tri-factorization to prioritize genetic loci and uncover associations contributing to complex trait heritability.


Key Features:

  • Integration of Omics Data: Merges genotype data with other omics datasets to jointly analyze relationships among genotypes, phenotypes, variant effects, and gene networks.
  • Corrected Non-Negative Matrix Tri-Factorization: Applies corrected non-negative matrix tri-factorization as the core algorithm for multi-view data factorization and relationship discovery.
  • Clustering Techniques: Employs clustering approaches to assess interconnectedness between genotypes, phenotypes, variant effects, and gene networks.
  • Identification of Loci–Trait Associations: Evaluates the damaging impact of genetic variants and their network interactions to identify associations between specific loci and traits.
  • Prioritization of Weak Loci: Enhances prioritization of weak loci that are often overlooked by standard regression models used in GWAS, addressing missing heritability.
  • Replication and Validation: Replicated 129 genes reported in global GWAS and supported 226 out of 265 genes (85%), including novel associations such as NLGN1, lipid metabolism regulator DAB1, and pleiotropic gene CARM1.
  • Population Structure Adjustment: Accounts for strong population structure by considering individuals' ancestry to ensure robustness across diverse cohorts.
  • Flexibility and Adaptability: Allows incorporation of various omics data sources to tailor analyses to specific research needs.

Scientific Applications:

  • Personalized Medicine: Provides integrative insights into complex disease mechanisms to inform candidate gene prioritization for precision medicine studies.
  • GWAS Augmentation and Missing Heritability: Complements standard GWAS regression by capturing weak-effect loci and network-mediated associations contributing to missing heritability.
  • Lipid Trait Genetics: Identifies regulators and pleiotropic genes relevant to lipid metabolism, exemplified by findings including DAB1 and CARM1.
  • Cross-Population Genetic Analysis: Enables analyses that remain robust across diverse cohorts by adjusting for ancestry and population structure.

Methodology:

Uses corrected non-negative matrix tri-factorization and clustering to integrate genotype, phenotype, variant-effect, and gene-network data, evaluates damaging impacts of variants and network interactions, and adjusts for population structure by considering individual ancestry.

Topics

Details

License:
CPL-1.0
Maturity:
Mature
Cost:
Free of charge
Tool Type:
library
Operating Systems:
Linux, Windows, Mac
Programming Languages:
R
Added:
8/9/2019
Last Updated:
11/24/2024

Operations

Publications

Leal LG, David A, Jarvelin M, Sebert S, Männikkö M, Karhunen V, Seaby E, Hoggart C, Sternberg MJE. Identification of disease-associated loci using machine learning for genotype and network data integration. Bioinformatics. 2019;35(24):5182-5190. doi:10.1093/bioinformatics/btz310. PMID:31070705. PMCID:PMC6954643.

PMID: 31070705
PMCID: PMC6954643
Funding: - Wellcome Trust: WT/104955/Z/14/Z - European Union’s Horizon 2020: 668303, NFBC1966 - Academy of Finland: 104781, 1114194, 120315, 129269, 24300796 - University of Oulu: 75617 - National Heart, Lung and Blood Institute: 1RL1MH083268-01, 5R01HL087679-02 - The National Institute of Mental Health: 5R01MH63706 - Medical Research Council: MR/M013138/1 - DynaHEALTH: H2020-633595 - National Institute of General Medical Sciences: U01-HG-004610 - University of Washington: U01-HG-004608 - Marshfield Clinic Research Foundation and Vanderbilt University Medical Center: U01-HG-04599 - Mayo Clinic: U01HG004609 - Northwestern University: U01-HG-04603 - Administrative Coordinating Center: U01HG004438 - Center for Inherited Disease Research: U01HG004424

Documentation

Links