Contrail
Contrail assembles large genomes from short-read second-generation sequencing data using Hadoop-based distributed computing to enable scalable de novo genome assembly in cloud environments.
Key Features:
- Hadoop-Based Infrastructure: Utilizes Hadoop for scalable distributed processing of large sequencing datasets.
- Cloud Compatibility: Operates in cloud environments to scale compute and storage for large-scale genome assembly.
- Short-Read Assembly: Specializes in assembling genomes from short-read data produced by second-generation sequencing technologies.
Scientific Applications:
- De novo assembly of large genomes: Produces genome assemblies from short reads for organisms with large or complex genomes.
- Genomic research on genetic diversity, evolution, and complex traits: Supports studies in genetic diversity, evolutionary biology, and complex trait analysis by providing assembled genomes from short-read data.
Methodology:
Processes different types of data produced by second-generation sequencers and accommodates changes in read lengths and sequencing technologies; incorporates assembly algorithms tailored for short-read assembly to address sequence coverage and error rates; uses Hadoop-based distributed computing in cloud environments to manage and process large datasets; and provides recommendations on sequencing strategies to enable high-quality genome assemblies.
Topics
Details
- Maturity:
- Mature
- Tool Type:
- command-line tool
- Operating Systems:
- Linux
- Added:
- 1/13/2017
- Last Updated:
- 11/25/2024
Operations
Publications
Schatz MC, Delcher AL, Salzberg SL. Assembly of large genomes using second-generation sequencing. Genome Research. 2010;20(9):1165-1173. doi:10.1101/gr.101360.109. PMID:20508146. PMCID:PMC2928494.
Documentation
Downloads
- Source codehttps://github.com/ndimiduk/contrail-bio