GeneGrouper
GeneGrouper identifies and clusters gene clusters across diverse genomic datasets to enable population-level analysis of genetic variation and the detection of conserved and variable gene architectures.
Key Features:
- Density-based clustering: Employs a density-based clustering method to group gene clusters based on similarity.
- Binning by similarity: Categorizes gene clusters into bins according to their pairwise similarity.
- Cross-taxa detection: Detects gene clusters across phylogenetically proximate and distant taxa.
- Sensitivity to small differences: Identifies conserved gene clusters that differ by only a few genes, including cases resulting in pseudogenes or regulatory interruptions.
- Distant homologs and mobile elements: Clusters both distant homologous gene clusters and variable gene clusters found in integrative and conjugative elements.
- Population-level characterization: Provides population-level characterization of genetic variation within similar gene clusters.
- Benchmark performance: Demonstrated high recall and precision in benchmarks on a mixed-taxa dataset of 435 genomes.
- Specific cluster identification: Detected the 23-gene Salmonella enterica LT2 Pdu gene cluster and the four-gene Pseudomonas aeruginosa PAO1 Mex gene cluster in benchmark datasets.
- Discovery and experimental linkage: Uncovered a frequently occurring pduN pseudogene in 1130 S. enterica genomes and linked frameshift disruption of pduN to impaired microcompartment formation in vivo.
Scientific Applications:
- Comparative genomics: Enables identification and comparison of gene cluster conservation and variation across genomes.
- Pseudogene and regulatory variation detection: Facilitates detection of pseudogenes and regulatory disruptions within conserved clusters.
- Evolutionary and population studies: Supports population-level studies of genetic diversity and evolutionary dynamics of gene clusters.
- Study of mobile genetic elements: Enables analysis of variable gene clusters in integrative and conjugative elements.
- Phenotype association: Assists in linking cluster variation to phenotypic effects, as exemplified by pduN frameshift effects on microcompartment formation.
Methodology:
Applies a density-based clustering algorithm to categorize gene clusters into similarity-based bins.
Topics
Details
- License:
- MIT
- Tool Type:
- command-line tool
- Programming Languages:
- Python, R
- Added:
- 9/20/2021
- Last Updated:
- 9/20/2021
Operations
Publications
McFarland AG, Kennedy NW, Mills CE, Tullman-Ercek D, Huttenhower C, Hartmann EM. Density-based binning of gene clusters to infer function or evolutionary history using GeneGrouper. Unknown Journal. 2021. doi:10.1101/2021.05.27.446007.
Links
Issue tracker
https://github.com/agmcfarland/GeneGrouper/issues