FSBC

FSBC performs fast string-based clustering of HT-SELEX oligonucleotide sequences to identify over-represented strings of varying lengths as potential aptamer target binding regions.


Key Features:

  • Single-round cluster estimation: Performs cluster estimation from a single (typically final) round of HT-SELEX data to identify sequence groups that include potential aptamers.
  • Over-represented string detection: Focuses on identifying over-represented strings as candidate target binding regions to reduce the search space.
  • Variable-length motif search: Searches for over-represented strings across multiple lengths to account for variability in target binding region length.
  • Robustness to nucleobase imbalance: Accommodates imbalanced nucleobase composition that often results from the aptamer selection process.
  • String-based clustering approach: Contrasts with conventional methods that rely on full-length sequence similarity or limited motif lengths by using string-focused clustering.
  • Large-scale performance: On HT-SELEX datasets exceeding 15 million oligonucleotide sequences, achieved the highest clustering accuracy and ranked second in calculation speed compared with FASTAptamer, AptaCluster, APTANI, and AptaTRACE.

Scientific Applications:

  • Aptamer candidate identification from HT-SELEX: Identifies sequence groups containing potential aptamers from HT-SELEX experiments.
  • Large-scale HT-SELEX analysis: Enables clustering and candidate selection on large HT-SELEX datasets, including analyses involving over 15 million oligonucleotides.
  • Detection of variable binding regions: Facilitates discovery of target binding regions with variable lengths across different molecules or binding styles.
  • Mitigation of selection bias effects: Addresses biases such as imbalanced nucleobase composition introduced during the selection process.

Methodology:

FSBC searches HT-SELEX sequences for over-represented strings across multiple lengths, reduces the search space by using those strings as candidate binding regions, and performs cluster estimation from a single (typically final) round of HT-SELEX data; comparative evaluation was performed against FASTAptamer, AptaCluster, APTANI, and AptaTRACE on datasets exceeding 15 million oligonucleotides.

Topics

Details

Tool Type:
library
Programming Languages:
R, Python
Added:
1/18/2021
Last Updated:
3/11/2021

Operations

Publications

Kato S, Ono T, Minagawa H, Horii K, Shiratori I, Waga I, Ito K, Aoki T. FSBC: fast string-based clustering for HT-SELEX data. BMC Bioinformatics. 2020;21(1). doi:10.1186/s12859-020-03607-1. PMID:32580745. PMCID:PMC7313139.