pclouds

pclouds identifies repeat structures in large eukaryotic genomes by calculating exact oligonucleotide counts to evaluate de novo repeat content.


Key Features:

  • Oligonucleotide Counting Approach: Calculates exact counts for any length oligonucleotide across large genomes without relying on sequence alignment or similarity searches.
  • P-cloud Construction: Builds oligonucleotide excess probability clouds ("P-clouds") that cluster related oligonucleotides occurring more frequently than expected by chance to indicate repetitive regions.
  • Efficiency and Speed: Analyzes whole genomes such as the human genome on a single desktop in less than half a day, reported as at least an order-of-magnitude (~10×) speed improvement over existing methods.
  • Mapping and Identification: Maps P-clouds back onto the genome and uses a sliding window approach to identify regions of high P-cloud density as repetitive regions.
  • Comprehensive Detection: Detects known repeat elements and additional repetitive sequences including gene families, pseudogenes, segmental duplicons, and transposable elements (TEs), with sensitivity to small fragments (~25 bp) beyond conventional methods like RepeatMasker.
  • Novel Element Identification: Develops element-specific P-clouds (ESPs) to identify novel repetitive elements such as Alu and MIR SINE sequences.

Scientific Applications:

  • De Novo Repeat Structure Identification: Identifies repeat structure in newly sequenced genomes without prior repeat libraries.
  • Transposable Element Analysis: Detects transposable elements and low-complexity repeats by locating clusters of high-abundance oligonucleotides.
  • Genome Annotation: Predicts additional repeat-derived content, informing probabilistic genome annotation and the assessment of unannotated repetitive sequences in genomes such as human.

Methodology:

Compute exact counts for any-length oligonucleotides; construct oligonucleotide excess probability clouds (P-clouds); map P-clouds onto the genome and apply a sliding window to identify high-density repetitive regions; generate element-specific P-clouds (ESPs) for novel Alu and MIR SINE identification.

Topics

Details

License:
Not licensed
Cost:
Free of charge
Tool Type:
command-line tool
Operating Systems:
Linux, Mac
Added:
3/23/2022
Last Updated:
3/23/2022

Operations

Publications

Gu W, Castoe TA, Hedges DJ, Batzer MA, Pollock DD. Identification of repeat structure in large genomes using repeat probability clouds. Analytical Biochemistry. 2008;380(1):77-83. doi:10.1016/j.ab.2008.05.015. PMID:18541131. PMCID:PMC2533575.

de Koning APJ, Gu W, Castoe TA, Batzer MA, Pollock DD. Repetitive Elements May Comprise Over Two-Thirds of the Human Genome. PLoS Genetics. 2011;7(12):e1002384. doi:10.1371/journal.pgen.1002384. PMID:22144907. PMCID:PMC3228813.

Documentation