cNMTF
cNMTF integrates genotype, phenotype, variant-effect, and gene-network data using corrected non-negative matrix tri-factorization to prioritize genetic loci and uncover associations contributing to complex trait heritability.
Key Features:
- Integration of Omics Data: Merges genotype data with other omics datasets to jointly analyze relationships among genotypes, phenotypes, variant effects, and gene networks.
- Corrected Non-Negative Matrix Tri-Factorization: Applies corrected non-negative matrix tri-factorization as the core algorithm for multi-view data factorization and relationship discovery.
- Clustering Techniques: Employs clustering approaches to assess interconnectedness between genotypes, phenotypes, variant effects, and gene networks.
- Identification of Loci–Trait Associations: Evaluates the damaging impact of genetic variants and their network interactions to identify associations between specific loci and traits.
- Prioritization of Weak Loci: Enhances prioritization of weak loci that are often overlooked by standard regression models used in GWAS, addressing missing heritability.
- Replication and Validation: Replicated 129 genes reported in global GWAS and supported 226 out of 265 genes (85%), including novel associations such as NLGN1, lipid metabolism regulator DAB1, and pleiotropic gene CARM1.
- Population Structure Adjustment: Accounts for strong population structure by considering individuals' ancestry to ensure robustness across diverse cohorts.
- Flexibility and Adaptability: Allows incorporation of various omics data sources to tailor analyses to specific research needs.
Scientific Applications:
- Personalized Medicine: Provides integrative insights into complex disease mechanisms to inform candidate gene prioritization for precision medicine studies.
- GWAS Augmentation and Missing Heritability: Complements standard GWAS regression by capturing weak-effect loci and network-mediated associations contributing to missing heritability.
- Lipid Trait Genetics: Identifies regulators and pleiotropic genes relevant to lipid metabolism, exemplified by findings including DAB1 and CARM1.
- Cross-Population Genetic Analysis: Enables analyses that remain robust across diverse cohorts by adjusting for ancestry and population structure.
Methodology:
Uses corrected non-negative matrix tri-factorization and clustering to integrate genotype, phenotype, variant-effect, and gene-network data, evaluates damaging impacts of variants and network interactions, and adjusts for population structure by considering individual ancestry.
Topics
Details
- License:
- CPL-1.0
- Maturity:
- Mature
- Cost:
- Free of charge
- Tool Type:
- library
- Operating Systems:
- Linux, Windows, Mac
- Programming Languages:
- R
- Added:
- 8/9/2019
- Last Updated:
- 11/24/2024
Operations
Publications
Leal LG, David A, Jarvelin M, Sebert S, Männikkö M, Karhunen V, Seaby E, Hoggart C, Sternberg MJE. Identification of disease-associated loci using machine learning for genotype and network data integration. Bioinformatics. 2019;35(24):5182-5190. doi:10.1093/bioinformatics/btz310. PMID:31070705. PMCID:PMC6954643.