CleanBSequences

CleanBSequences removes adaptor and cloning-vector-derived sequences from PCR-amplified biological sequence datasets to automate curation for AFLP, cDNA-AFLP, and MSAP-based genomic, transcriptomic, and epigenomic studies.


Key Features:

  • Alignment-based trimming: Aligns forward and/or reverse primers or cloning vector ends to target sequences to identify and remove adaptor and undesired fragments.
  • Adaptor removal for PCR amplification: Detects and excises adaptor sequences used for PCR amplification from sequence reads.
  • Subsequence retention: Retains necessary target subsequences after trimming to preserve biological signal.
  • High-throughput processing: Processes large datasets rapidly and accurately for omic-scale studies.
  • R package implementation: Implemented as an R package for integration into R-based bioinformatics workflows.
  • Error reduction: Automates curation to reduce errors associated with manual sequence trimming.

Scientific Applications:

  • Molecular marker preprocessing: Cleans sequence data from AFLP, cDNA-AFLP, and MSAP experiments prior to downstream analysis.
  • Omic data curation: Prepares genomic, transcriptomic, and epigenomic datasets by removing adaptor and cloning-vector contamination.
  • Sequence analysis preparation: Streamlines downstream sequence analysis workflows by standardizing trimmed sequences.

Methodology:

Aligns forward and/or reverse primers or cloning vector ends to target sequences and removes matching adaptor and undesired fragments while retaining target subsequences.

Topics

Details

License:
GPL-3.0
Tool Type:
library
Programming Languages:
R
Added:
1/18/2021
Last Updated:
2/11/2021

Operations

Publications

Pozzi FI, Green GY, Barbona IG, Rodríguez GR, Felitti SA. CleanBSequences: an efficient curator of biological sequences in R. Molecular Genetics and Genomics. 2020;295(4):837-841. doi:10.1007/s00438-020-01671-z. PMID:32300860.

PMID: 32300860
Funding: - Consejo Nacional de Investigaciones Científicas y Técnicas: PIP11220090100613, PUE22920160100043CO - Fondo para la Investigación Científica y Tecnológica: PICT20121321