hapFabia
hapFabia identifies very short DNA segments that are identical by descent (IBD) in large sequencing and genotype datasets by exploiting rare variants to increase detection resolution.
Key Features:
- Biclustering Approach: Employs a biclustering algorithm to detect very short IBD segments characterized by rare variants.
- Variant Types and Input Data: Uses rare variants and processes genotype data from next-generation sequencing and DNA microarray platforms, accepting Variant Call Format (VCF) files.
- Sparse Matrix Implementation: Utilizes sparse matrix operations to efficiently process large genomic datasets.
- Detection Performance: Demonstrated superior performance on artificial and simulated datasets for detecting short IBD segments tagged by rare variants.
- Empirical Results: Detected 160,588 short IBD segments in 1000 Genomes Project chromosome 1 data with median length 23 kb and mean 24 kb, comprising approximately 752,000 SNVs (~39% of rare variants and 23.5% of all variants).
- Population Sharing and Ancient Matches: Identifies that most short IBD segments are shared among African populations with fewer in Europeans and Asians, and reports matches to Denisova and Neandertal genomes more frequently in Asians and Europeans and occasionally exclusively in Africans.
- Evolutionary Implications: Length distributions and sharing patterns indicate many identified IBD segments predate the out-of-Africa migration.
- Visualization: Provides visualization of haplotype clusters in various formats to represent identified IBD segments.
Scientific Applications:
- Phasing and Imputation: Improves phasing accuracy and imputation quality, particularly for low-coverage sequencing, by detecting short IBD segments.
- Association Studies: Increases the power of association studies by providing high-resolution identification of short IBD segments.
- Population Genetics: Reveals ancient genetic connections and evolutionary patterns across human populations through analysis of short IBD segments.
Methodology:
Applies biclustering on genotype matrices exploiting rare variants, processes VCF genotype data from next-generation sequencing and microarray sources using sparse matrix operations, and generates haplotype-cluster visualizations.
Topics
Collections
Details
- License:
- GPL-2.0
- Tool Type:
- command-line tool, library
- Operating Systems:
- Linux, Windows, Mac
- Programming Languages:
- R
- Added:
- 1/17/2017
- Last Updated:
- 1/10/2019
Operations
Data Inputs & Outputs
Publications
Hochreiter S. HapFABIA: Identification of very short segments of identity by descent characterized by rare variants in large sequencing data. Nucleic Acids Research. 2013;41(22):e202-e202. doi:10.1093/nar/gkt1013. PMID:24174545. PMCID:PMC3905877.