Sequali

Sequali performs comprehensive quality control on short- and long-read sequencing data by identifying adapter sequences, overrepresented sequences, and duplicates to support downstream analyses.


Key Features:

  • Adapter Search: Identifies and reports adapter sequences present in sequencing reads.
  • Overrepresented Sequence Analysis: Detects and characterizes overrepresented sequences to reveal biases or contaminants.
  • Duplication Analysis: Identifies duplicate reads to assess redundancy and potential PCR artifacts.
  • Support for Multiple Input Formats: Accepts FASTQ and uBAM file formats for both short- and long-read datasets.
  • High-performance Implementation: Implements performance optimizations using Python with C extensions to accelerate processing.

Scientific Applications:

  • Cross-platform sequencing QC: Performs quality control for datasets from short-read and long-read technologies, including Oxford Nanopore Technologies.
  • Variant calling preparation: Produces QC metrics and filtered reads to improve reliability of variant calling workflows.
  • Genome assembly preprocessing: Identifies contaminants and duplicates to improve accuracy of genome assembly.
  • Transcriptome profiling QC: Detects sequence-level artifacts that can affect transcriptome quantification and analysis.

Methodology:

Implemented in Python with C extensions to improve performance and reported to operate faster than comparable quality control programs for short- and long-read sequencing.

Topics

Details

Tool Type:
command-line tool
Added:
2/14/2025
Last Updated:
2/14/2025

Operations

Publications

Vorderman RHP. Sequali: efficient and comprehensive quality control of short- and long-read sequencing data. Bioinformatics Advances. 2024;5(1). doi:10.1093/bioadv/vbaf010. PMID:39927290. PMCID:PMC11802474.