GeneGrouper

GeneGrouper identifies and clusters gene clusters across diverse genomic datasets to enable population-level analysis of genetic variation and the detection of conserved and variable gene architectures.


Key Features:

  • Density-based clustering: Employs a density-based clustering method to group gene clusters based on similarity.
  • Binning by similarity: Categorizes gene clusters into bins according to their pairwise similarity.
  • Cross-taxa detection: Detects gene clusters across phylogenetically proximate and distant taxa.
  • Sensitivity to small differences: Identifies conserved gene clusters that differ by only a few genes, including cases resulting in pseudogenes or regulatory interruptions.
  • Distant homologs and mobile elements: Clusters both distant homologous gene clusters and variable gene clusters found in integrative and conjugative elements.
  • Population-level characterization: Provides population-level characterization of genetic variation within similar gene clusters.
  • Benchmark performance: Demonstrated high recall and precision in benchmarks on a mixed-taxa dataset of 435 genomes.
  • Specific cluster identification: Detected the 23-gene Salmonella enterica LT2 Pdu gene cluster and the four-gene Pseudomonas aeruginosa PAO1 Mex gene cluster in benchmark datasets.
  • Discovery and experimental linkage: Uncovered a frequently occurring pduN pseudogene in 1130 S. enterica genomes and linked frameshift disruption of pduN to impaired microcompartment formation in vivo.

Scientific Applications:

  • Comparative genomics: Enables identification and comparison of gene cluster conservation and variation across genomes.
  • Pseudogene and regulatory variation detection: Facilitates detection of pseudogenes and regulatory disruptions within conserved clusters.
  • Evolutionary and population studies: Supports population-level studies of genetic diversity and evolutionary dynamics of gene clusters.
  • Study of mobile genetic elements: Enables analysis of variable gene clusters in integrative and conjugative elements.
  • Phenotype association: Assists in linking cluster variation to phenotypic effects, as exemplified by pduN frameshift effects on microcompartment formation.

Methodology:

Applies a density-based clustering algorithm to categorize gene clusters into similarity-based bins.

Topics

Details

License:
MIT
Tool Type:
command-line tool
Programming Languages:
Python, R
Added:
9/20/2021
Last Updated:
9/20/2021

Operations

Publications

McFarland AG, Kennedy NW, Mills CE, Tullman-Ercek D, Huttenhower C, Hartmann EM. Density-based binning of gene clusters to infer function or evolutionary history using GeneGrouper. Unknown Journal. 2021. doi:10.1101/2021.05.27.446007.

Links