FSBC
FSBC performs fast string-based clustering of HT-SELEX oligonucleotide sequences to identify over-represented strings of varying lengths as potential aptamer target binding regions.
Key Features:
- Single-round cluster estimation: Performs cluster estimation from a single (typically final) round of HT-SELEX data to identify sequence groups that include potential aptamers.
- Over-represented string detection: Focuses on identifying over-represented strings as candidate target binding regions to reduce the search space.
- Variable-length motif search: Searches for over-represented strings across multiple lengths to account for variability in target binding region length.
- Robustness to nucleobase imbalance: Accommodates imbalanced nucleobase composition that often results from the aptamer selection process.
- String-based clustering approach: Contrasts with conventional methods that rely on full-length sequence similarity or limited motif lengths by using string-focused clustering.
- Large-scale performance: On HT-SELEX datasets exceeding 15 million oligonucleotide sequences, achieved the highest clustering accuracy and ranked second in calculation speed compared with FASTAptamer, AptaCluster, APTANI, and AptaTRACE.
Scientific Applications:
- Aptamer candidate identification from HT-SELEX: Identifies sequence groups containing potential aptamers from HT-SELEX experiments.
- Large-scale HT-SELEX analysis: Enables clustering and candidate selection on large HT-SELEX datasets, including analyses involving over 15 million oligonucleotides.
- Detection of variable binding regions: Facilitates discovery of target binding regions with variable lengths across different molecules or binding styles.
- Mitigation of selection bias effects: Addresses biases such as imbalanced nucleobase composition introduced during the selection process.
Methodology:
FSBC searches HT-SELEX sequences for over-represented strings across multiple lengths, reduces the search space by using those strings as candidate binding regions, and performs cluster estimation from a single (typically final) round of HT-SELEX data; comparative evaluation was performed against FASTAptamer, AptaCluster, APTANI, and AptaTRACE on datasets exceeding 15 million oligonucleotides.
Topics
Details
- Tool Type:
- library
- Programming Languages:
- R, Python
- Added:
- 1/18/2021
- Last Updated:
- 3/11/2021
Operations
Publications
Kato S, Ono T, Minagawa H, Horii K, Shiratori I, Waga I, Ito K, Aoki T. FSBC: fast string-based clustering for HT-SELEX data. BMC Bioinformatics. 2020;21(1). doi:10.1186/s12859-020-03607-1. PMID:32580745. PMCID:PMC7313139.