Clean reads

Clean reads processes next-generation sequencing (NGS) data from Sanger, 454, Illumina, and SOLiD platforms to trim adaptors and low-quality regions and produce cleaned reads for downstream analyses such as transcriptome assembly, read mapping, and SNP discovery.


Key Features:

  • Read cleaning: Trims low-quality regions and removes adaptors, vectors, and unwanted sequences defined by regular expressions.
  • Quality filtering: Removes reads that do not meet predefined quality criteria.
  • Parallel processing pipeline (ngs_backbone): Executes cleaning and downstream steps within a parallel pipeline framework to scale to large NGS datasets.
  • Platform support: Accepts input from Sanger EST reads, 454, Illumina, and SOLiD sequencing technologies.
  • Transcriptome assembly and annotation support: Prepares cleaned reads for construction and functional annotation of transcriptomes.
  • Read mapping: Aligns cleaned reads to reference genomes or transcriptomes, including mapping to the SGN tomato transcriptome.
  • SNP calling and selection: Identifies single nucleotide polymorphisms and reports predicted single nucleotide variations (SNVs).

Scientific Applications:

  • SNP discovery in non-model species: Enables identification and selection of SNPs from diverse organisms by processing data from multiple sequencing platforms.
  • Tomato (Solanum lycopersicum) case study: Processed Sanger EST reads and Illumina data, mapped to the SGN tomato transcriptome, predicting 23,360 SNVs with 85% experimental validation.

Methodology:

Trimming of low-quality regions, adaptor and vector removal using regular expressions, quality-based read filtering, parallel execution within the ngs_backbone pipeline, read mapping to reference genomes/transcriptomes, and SNP calling; integration with high-throughput genotyping data and NGS data analysis.

Topics

Details

Maturity:
Legacy
Tool Type:
command-line tool
Operating Systems:
Linux
Programming Languages:
Python
Added:
1/13/2017
Last Updated:
11/25/2024

Operations

Publications

Blanca JM, Pascual L, Ziarsolo P, Nuez F, Cañizares J. ngs_backbone: a pipeline for read cleaning, mapping and SNP calling using Next Generation Sequence. BMC Genomics. 2011;12(1). doi:10.1186/1471-2164-12-285. PMID:21635747. PMCID:PMC3124440.