Neptune

Neptune identifies differentially abundant genomic loci in bacterial populations using k-mer-based matching and probabilistic comparative-genomics methods.


Key Features:

  • Exact k-mer matching with mismatch accommodation: Uses exact k-mer matching while allowing mismatches to locate genomic signatures with high precision.
  • Probabilistic loci discovery: Employs probabilistic models rather than heuristic strategies to identify loci that are common in target groups and rare in non-target groups.
  • Parallel computing for efficiency: Leverages parallel computing to handle large datasets and reduce processing time.
  • Assembly-based locus extraction without multiple sequence alignments: Identifies and extracts relevant loci from draft genome assemblies without requiring multiple sequence alignments or other computationally intensive comparative analyses.
  • Sensitivity and specificity evaluation: Demonstrates rapid identification of regions that are both sensitive and specific based on testing with simulated and real datasets.
  • Differential abundance detection: Detects sequences significantly represented in target groups while absent or rare in non-target groups.

Scientific Applications:

  • Comparative bacterial genomics: Discovery of loci that distinguish bacterial populations and lineages.
  • Pathogenomics: Identification of pathogen-associated genomic signatures relevant to virulence and epidemiology.
  • Trait-associated locus discovery: Detection of loci associated with specific traits such as antibiotic resistance or virulence factors.
  • Study of bacterial evolution: Analysis of differentially abundant genomic content to infer evolutionary relationships and lineage-specific variation.

Methodology:

Neptune applies exact k-mer matching with mismatch accommodation, probabilistic models for loci discovery, and parallel computing to identify and extract loci from draft genome assemblies without requiring multiple sequence alignments.

Topics

Details

Tool Type:
command-line tool
Operating Systems:
Linux, Mac
Programming Languages:
Shell, Python
Added:
5/21/2018
Last Updated:
11/25/2024

Operations

Data Inputs & Outputs

Nucleic acid sequence analysis

Publications

Marinier E, Zaheer R, Berry C, Weedmark KA, Domaratzki M, Mabon P, Knox NC, Reimer AR, Graham MR, Chui L, Patterson-Fortin L, Zhang J, Pagotto F, Farber J, Mahony J, Seyer K, Bekal S, Tremblay C, Isaac-Renton J, Prystajecky N, Chen J, Slade P, Van Domselaar G. Neptune: a bioinformatics tool for rapid discovery of genomic variation in bacterial populations. Nucleic Acids Research. 2017;45(18):e159-e159. doi:10.1093/nar/gkx702. PMID:29048594. PMCID:PMC5737611.

Documentation