GET_HOMOLOGUES
GET_HOMOLOGUES performs pan-genome clustering and comparative analyses of homologous gene families and, with the GET_HOMOLOGUES-EST extension, analyzes pantranscriptomes from RNA-seq and transcriptomic data to classify core and accessory sequences and estimate pangenome properties.
Key Features:
- Homology clustering algorithms: Supports clustering with bidirectional best-hit, COGtriangles, and OrthoMCL algorithms.
- Domain composition scanning: Integrates HMMER3-based domain composition screening to inform cluster stringency.
- Alignment and synteny filters: Allows pairwise alignment coverage cutoffs and selection of syntenic genes to refine homologous clusters.
- Consensus clustering: Produces robust homologous gene families via consensus clustering from combinations of algorithms and filters.
- Pangenome and core genome construction: Constructs, interrogates, and generates graphical representations of core genome and pangenome sets.
- Statistical pangenome models: Uses exponential and binomial mixture models to estimate theoretical pangenome and core genome sizes.
- Pangenome trees and comparative analyses: Computes pangenome trees and performs comparative genomics tasks including identification of lineage-specific genes and gene family expansions.
- Pantranscriptome support (GET_HOMOLOGUES-EST): Leverages transcriptomic and RNA-seq data to sample genic repertoires for species with large or highly heterozygous genomes.
- Presence–absence and ANI matrices: Computes presence–absence pangenome matrices and compiles Average Nucleotide Identity (ANI) matrices summarizing intra-species variation.
- Pan-genome growth simulation and private gene detection: Simulates pangenome growth from sequence clusters and identifies private genes present only in some genotypes.
- Parallelization: Parallelizes computationally intensive tasks for multiprocessor workstations and computer clusters.
- Empirical benchmarking: Has been benchmarked on datasets including 50 Streptococcus genomes (OMA annotations), 19 Arabidopsis thaliana ecotypes, and transcripts from 16 Hordeum vulgare genotypes.
- Accessory locus characterization: Enables analysis showing accessory loci tend to have lower expression, higher non-synonymous substitution rates, and enrichment for transposon components, disease resistance genes, and presence/absence-associated protein domains.
Scientific Applications:
- Microbial pan-genomics: Comparative analysis and reconstruction of core and accessory gene sets across microbial genomes.
- Plant pantranscriptomics and pangene analysis: Sampling and classification of genic repertoires in species with large or heterozygous genomes using RNA-seq-derived transcripts.
- Pangenome size estimation: Estimation of theoretical pangenome and core genome sizes using statistical mixture models.
- Phylogenetics and population structure: Reconstruction of pangenome trees and summarization of intra-species variation via ANI matrices.
- Presence/absence variation studies: Detection and characterization of accessory sequences, private genes, and loci associated with transposons and disease resistance.
- Comparative genomics: Identification of lineage-specific genes, gene family expansions, and differences in expression and substitution patterns between core and accessory genes.
Methodology:
Clusters homologous genes using bidirectional best-hit, COGtriangles, and OrthoMCL; applies HMMER3 domain scans, pairwise alignment coverage cutoffs, and synteny filters; produces consensus clusters; fits exponential and binomial mixture models for pangenome size estimation; computes pangenome trees, presence–absence matrices, and ANI matrices; uses transcriptomic/RNA-seq data for GET_HOMOLOGUES-EST benchmarking on specified datasets and parallelizes tasks for multiprocessor execution.
Topics
Details
- License:
- Other
- Maturity:
- Mature
- Cost:
- Free of charge (with restrictions)
- Tool Type:
- command-line tool
- Operating Systems:
- Linux, Mac
- Added:
- 10/23/2023
- Last Updated:
- 11/24/2024
Operations
Publications
Contreras-Moreira B, Vinuesa P. GET_HOMOLOGUES, a Versatile Software Package for Scalable and Robust Microbial Pangenome Analysis. Applied and Environmental Microbiology. 2013;79(24):7696-7701. doi:10.1128/aem.02411-13. PMID:24096415. PMCID:PMC3837814.
Contreras-Moreira B, Cantalapiedra CP, García-Pereira MJ, Gordon SP, Vogel JP, Igartua E, Casas AM, Vinuesa P. Analysis of Plant Pan-Genomes and Transcriptomes with GET_HOMOLOGUES-EST, a Clustering Solution for Sequences of the Same Species. Frontiers in Plant Science. 2017;8. doi:10.3389/fpls.2017.00184. PMID:28261241. PMCID:PMC5306281.
Vinuesa P, Contreras-Moreira B. Robust Identification of Orthologues and Paralogues for Microbial Pan-Genomics Using GET_HOMOLOGUES: A Case Study of pIncA/C Plasmids. Methods in Molecular Biology. 2015. doi:10.1007/978-1-4939-1720-4_14. PMID:25343868.
Contreras-Moreira B, del Río ÁR, Cantalapiedra CP, Sancho R, Vinuesa P. Pangenome Analysis of Plant Transcripts and Coding Sequences. Methods in Molecular Biology. 2022. doi:10.1007/978-1-0716-2429-6_9. PMID:35818004.
Documentation
Downloads
- Container filehttps://hub.docker.com/r/csicunam/get_homologues
- Downloads pagehttps://github.com/eead-csic-compbio/get_homologues/releases
- Otherhttps://bioconda.github.io/recipes/get_homologues/README.htmlConda package
- Source codehttps://github.com/eead-csic-compbio/get_homologues