GET_HOMOLOGUES

GET_HOMOLOGUES performs pan-genome clustering and comparative analyses of homologous gene families and, with the GET_HOMOLOGUES-EST extension, analyzes pantranscriptomes from RNA-seq and transcriptomic data to classify core and accessory sequences and estimate pangenome properties.


Key Features:

  • Homology clustering algorithms: Supports clustering with bidirectional best-hit, COGtriangles, and OrthoMCL algorithms.
  • Domain composition scanning: Integrates HMMER3-based domain composition screening to inform cluster stringency.
  • Alignment and synteny filters: Allows pairwise alignment coverage cutoffs and selection of syntenic genes to refine homologous clusters.
  • Consensus clustering: Produces robust homologous gene families via consensus clustering from combinations of algorithms and filters.
  • Pangenome and core genome construction: Constructs, interrogates, and generates graphical representations of core genome and pangenome sets.
  • Statistical pangenome models: Uses exponential and binomial mixture models to estimate theoretical pangenome and core genome sizes.
  • Pangenome trees and comparative analyses: Computes pangenome trees and performs comparative genomics tasks including identification of lineage-specific genes and gene family expansions.
  • Pantranscriptome support (GET_HOMOLOGUES-EST): Leverages transcriptomic and RNA-seq data to sample genic repertoires for species with large or highly heterozygous genomes.
  • Presence–absence and ANI matrices: Computes presence–absence pangenome matrices and compiles Average Nucleotide Identity (ANI) matrices summarizing intra-species variation.
  • Pan-genome growth simulation and private gene detection: Simulates pangenome growth from sequence clusters and identifies private genes present only in some genotypes.
  • Parallelization: Parallelizes computationally intensive tasks for multiprocessor workstations and computer clusters.
  • Empirical benchmarking: Has been benchmarked on datasets including 50 Streptococcus genomes (OMA annotations), 19 Arabidopsis thaliana ecotypes, and transcripts from 16 Hordeum vulgare genotypes.
  • Accessory locus characterization: Enables analysis showing accessory loci tend to have lower expression, higher non-synonymous substitution rates, and enrichment for transposon components, disease resistance genes, and presence/absence-associated protein domains.

Scientific Applications:

  • Microbial pan-genomics: Comparative analysis and reconstruction of core and accessory gene sets across microbial genomes.
  • Plant pantranscriptomics and pangene analysis: Sampling and classification of genic repertoires in species with large or heterozygous genomes using RNA-seq-derived transcripts.
  • Pangenome size estimation: Estimation of theoretical pangenome and core genome sizes using statistical mixture models.
  • Phylogenetics and population structure: Reconstruction of pangenome trees and summarization of intra-species variation via ANI matrices.
  • Presence/absence variation studies: Detection and characterization of accessory sequences, private genes, and loci associated with transposons and disease resistance.
  • Comparative genomics: Identification of lineage-specific genes, gene family expansions, and differences in expression and substitution patterns between core and accessory genes.

Methodology:

Clusters homologous genes using bidirectional best-hit, COGtriangles, and OrthoMCL; applies HMMER3 domain scans, pairwise alignment coverage cutoffs, and synteny filters; produces consensus clusters; fits exponential and binomial mixture models for pangenome size estimation; computes pangenome trees, presence–absence matrices, and ANI matrices; uses transcriptomic/RNA-seq data for GET_HOMOLOGUES-EST benchmarking on specified datasets and parallelizes tasks for multiprocessor execution.

Topics

Details

License:
Other
Maturity:
Mature
Cost:
Free of charge (with restrictions)
Tool Type:
command-line tool
Operating Systems:
Linux, Mac
Added:
10/23/2023
Last Updated:
11/24/2024

Operations

Publications

Contreras-Moreira B, Vinuesa P. GET_HOMOLOGUES, a Versatile Software Package for Scalable and Robust Microbial Pangenome Analysis. Applied and Environmental Microbiology. 2013;79(24):7696-7701. doi:10.1128/aem.02411-13. PMID:24096415. PMCID:PMC3837814.

Contreras-Moreira B, Cantalapiedra CP, García-Pereira MJ, Gordon SP, Vogel JP, Igartua E, Casas AM, Vinuesa P. Analysis of Plant Pan-Genomes and Transcriptomes with GET_HOMOLOGUES-EST, a Clustering Solution for Sequences of the Same Species. Frontiers in Plant Science. 2017;8. doi:10.3389/fpls.2017.00184. PMID:28261241. PMCID:PMC5306281.

PMID: 28261241
PMCID: PMC5306281
Funding: - Fundación Bancaria Caixa d’Estalvis i Pensions de Barcelona: GA-LC-059-2011 - Ministerio de Economía y Competitividad: AGL2013-48756-R, CSIC13-4E-2490 - Consejo Nacional de Ciencia y Tecnología: 179133 - Dirección General de Asuntos del Personal Académico, Universidad Nacional Autónoma de México: IN211814 - U.S. Department of Energy: DE-AC02-05CH11231

Vinuesa P, Contreras-Moreira B. Robust Identification of Orthologues and Paralogues for Microbial Pan-Genomics Using GET_HOMOLOGUES: A Case Study of pIncA/C Plasmids. Methods in Molecular Biology. 2015. doi:10.1007/978-1-4939-1720-4_14. PMID:25343868.

Contreras-Moreira B, del Río ÁR, Cantalapiedra CP, Sancho R, Vinuesa P. Pangenome Analysis of Plant Transcripts and Coding Sequences. Methods in Molecular Biology. 2022. doi:10.1007/978-1-0716-2429-6_9. PMID:35818004.

Documentation

Training material
http://eead-csic-compbio.github.io/get_homologues/plant_pangenome/protocol.html
The pangenome of a species is the sum of the genomes of its individuals. As coding sequences often represent only a small fraction of each genome, analyzing the pangene set can be a cost-effective strategy for plants with large genomes or highly heterozygous species. Here we describe a step-by-step protocol to analyze plant pangene sets with the software GET_HOMOLOGUES-EST. After a short introduction, where the main concepts are illustrated, the remaining sections cover the installation and typical operations required to analyze and annotate pantranscriptomes and gene sets of plants. The recipes include instructions on how to call core and accessory genes, how to compute a presence-absence pangenome matrix and how to identify and analyze private genes, present only in some genotypes. Downstream phylogenetic analyses are also discussed.
Training material
http://eead-csic-compbio.github.io/get_homologues/tutorial/pangenome_tutorial.html
This tutorial illustrates how to analyze pan-genomes using GET_HOMOLOGUES and GET_HOMOLOGUES-EST. After a short introduction, where the main concepts are illustrated, the remaining sections cover the installation and typical operations required to analyze and annotate genomes and transcriptomes from a pan-genome perspective, in which individuals or species contribute genetic material to a pool.

Downloads

Links