pclouds
pclouds identifies repeat structures in large eukaryotic genomes by calculating exact oligonucleotide counts to evaluate de novo repeat content.
Key Features:
- Oligonucleotide Counting Approach: Calculates exact counts for any length oligonucleotide across large genomes without relying on sequence alignment or similarity searches.
- P-cloud Construction: Builds oligonucleotide excess probability clouds ("P-clouds") that cluster related oligonucleotides occurring more frequently than expected by chance to indicate repetitive regions.
- Efficiency and Speed: Analyzes whole genomes such as the human genome on a single desktop in less than half a day, reported as at least an order-of-magnitude (~10×) speed improvement over existing methods.
- Mapping and Identification: Maps P-clouds back onto the genome and uses a sliding window approach to identify regions of high P-cloud density as repetitive regions.
- Comprehensive Detection: Detects known repeat elements and additional repetitive sequences including gene families, pseudogenes, segmental duplicons, and transposable elements (TEs), with sensitivity to small fragments (~25 bp) beyond conventional methods like RepeatMasker.
- Novel Element Identification: Develops element-specific P-clouds (ESPs) to identify novel repetitive elements such as Alu and MIR SINE sequences.
Scientific Applications:
- De Novo Repeat Structure Identification: Identifies repeat structure in newly sequenced genomes without prior repeat libraries.
- Transposable Element Analysis: Detects transposable elements and low-complexity repeats by locating clusters of high-abundance oligonucleotides.
- Genome Annotation: Predicts additional repeat-derived content, informing probabilistic genome annotation and the assessment of unannotated repetitive sequences in genomes such as human.
Methodology:
Compute exact counts for any-length oligonucleotides; construct oligonucleotide excess probability clouds (P-clouds); map P-clouds onto the genome and apply a sliding window to identify high-density repetitive regions; generate element-specific P-clouds (ESPs) for novel Alu and MIR SINE identification.
Topics
Details
- License:
- Not licensed
- Cost:
- Free of charge
- Tool Type:
- command-line tool
- Operating Systems:
- Linux, Mac
- Added:
- 3/23/2022
- Last Updated:
- 3/23/2022
Operations
Publications
Gu W, Castoe TA, Hedges DJ, Batzer MA, Pollock DD. Identification of repeat structure in large genomes using repeat probability clouds. Analytical Biochemistry. 2008;380(1):77-83. doi:10.1016/j.ab.2008.05.015. PMID:18541131. PMCID:PMC2533575.
de Koning APJ, Gu W, Castoe TA, Batzer MA, Pollock DD. Repetitive Elements May Comprise Over Two-Thirds of the Human Genome. PLoS Genetics. 2011;7(12):e1002384. doi:10.1371/journal.pgen.1002384. PMID:22144907. PMCID:PMC3228813.