Clean reads
Clean reads processes next-generation sequencing (NGS) data from Sanger, 454, Illumina, and SOLiD platforms to trim adaptors and low-quality regions and produce cleaned reads for downstream analyses such as transcriptome assembly, read mapping, and SNP discovery.
Key Features:
- Read cleaning: Trims low-quality regions and removes adaptors, vectors, and unwanted sequences defined by regular expressions.
- Quality filtering: Removes reads that do not meet predefined quality criteria.
- Parallel processing pipeline (ngs_backbone): Executes cleaning and downstream steps within a parallel pipeline framework to scale to large NGS datasets.
- Platform support: Accepts input from Sanger EST reads, 454, Illumina, and SOLiD sequencing technologies.
- Transcriptome assembly and annotation support: Prepares cleaned reads for construction and functional annotation of transcriptomes.
- Read mapping: Aligns cleaned reads to reference genomes or transcriptomes, including mapping to the SGN tomato transcriptome.
- SNP calling and selection: Identifies single nucleotide polymorphisms and reports predicted single nucleotide variations (SNVs).
Scientific Applications:
- SNP discovery in non-model species: Enables identification and selection of SNPs from diverse organisms by processing data from multiple sequencing platforms.
- Tomato (Solanum lycopersicum) case study: Processed Sanger EST reads and Illumina data, mapped to the SGN tomato transcriptome, predicting 23,360 SNVs with 85% experimental validation.
Methodology:
Trimming of low-quality regions, adaptor and vector removal using regular expressions, quality-based read filtering, parallel execution within the ngs_backbone pipeline, read mapping to reference genomes/transcriptomes, and SNP calling; integration with high-throughput genotyping data and NGS data analysis.
Topics
Details
- Maturity:
- Legacy
- Tool Type:
- command-line tool
- Operating Systems:
- Linux
- Programming Languages:
- Python
- Added:
- 1/13/2017
- Last Updated:
- 11/25/2024
Operations
Publications
Blanca JM, Pascual L, Ziarsolo P, Nuez F, Cañizares J. ngs_backbone: a pipeline for read cleaning, mapping and SNP calling using Next Generation Sequence. BMC Genomics. 2011;12(1). doi:10.1186/1471-2164-12-285. PMID:21635747. PMCID:PMC3124440.