CFSP
CFSP discovers frequent sequence patterns in nucleic acid datasets to identify conserved motifs associated with transcriptional regulation and functional genomic elements.
Key Features:
- Frequent Sequence Pattern Discovery: Identifies conserved subsequences in nucleic acid datasets using a Teiresias-like feature extraction approach.
- Gap-Tolerant Pattern Detection: Detects sequence motifs with long-distance correlations by allowing large gaps between subsequences.
- Mutation-Aware Pattern Matching: Incorporates single nucleotide variants and insertion–deletion mutations (indels) to enable fuzzy matching of sequence patterns.
- Predictive Feature Extraction: Extracts sequence patterns used for predictive modeling of functional genomic elements.
Scientific Applications:
- Regulatory Motif Discovery: Identifies conserved sequence motifs involved in transcriptional regulation and protein–DNA interactions.
- Non-Coding RNA Identification: Detects sequence features associated with miRNA, piRNA, and Sigma 54 promoter regions.
Methodology:
CFSP constructs arrays of frequent subsequences from ordered nucleic acid sequences and encodes mutation information, including single nucleotide variants and indels, to identify conserved sequence patterns with fuzzy matching.
Topics
Details
- License:
- MIT
- Programming Languages:
- Shell, C, C++
- Added:
- 1/18/2021
- Last Updated:
- 2/10/2021
Operations
Publications
Peng H. CFSP: a collaborative frequent sequence pattern discovery algorithm for nucleic acid sequence classification. PeerJ. 2020;8:e8965. doi:10.7717/peerj.8965. PMID:32341900. PMCID:PMC7179567.