Cascabel

Cascabel processes amplicon sequence data to analyze marker gene sequencing targets such as the rRNA operon (16S, 18S, ITS) and cytochrome c oxidase I (CO1) for assessment of microbial and environmental DNA (eDNA) communities.


Key Features:

  • Workflow engine: Implemented with Snakemake for workflow management.
  • Input data: Accepts raw sequence data as input.
  • Outputs: Produces an operational taxonomic unit (OTU) table and a representative sequence tree.
  • OTU generation: Provides multiple selectable OTU generating methods.
  • Methods integration: Integrates existing solutions and newly developed computational methods.
  • Scalability: Supports execution across computing environments from personal computers to high-performance computing servers.

Scientific Applications:

  • Marker gene sequencing: Analysis of single-gene marker sequences (16S, 18S, ITS, CO1) for taxonomic profiling.
  • Microbial community assessment: Characterization of microbial communities in diverse environmental samples.
  • eDNA surveys: Environmental DNA sequencing analyses targeting multicellular organisms.
  • Plant and animal microbiomes: Investigation of microbiomes associated with plants and animals.
  • Cost-effective diversity surveys: Single-gene sequencing alternative to whole-genome approaches for biodiversity assessment.

Methodology:

Implemented as a Snakemake workflow that accepts raw sequence data, integrates existing and newly developed computational methods, supports multiple OTU generating methods, and outputs an OTU table and a representative sequence tree.

Topics

Details

License:
GPL-3.0
Programming Languages:
Java, Python
Added:
1/9/2020
Last Updated:
12/10/2020

Operations

Publications

Asbun AA, Besseling MA, Balzano S, van Bleijswijk J, Witte H, Villanueva L, Engelmann JC. <i>Cascabel</i>: a flexible, scalable and easy-to-use amplicon sequence data analysis pipeline. Unknown Journal. 2019. doi:10.1101/809384.