CLUSTOM-CLOUD

CLUSTOM-CLOUD clusters 16S rRNA sequence reads into operational taxonomic units (OTUs) using an In-Memory Data Grid to enable scalable distributed processing for microbial diversity analysis.


Key Features:

  • In-Memory Data Grid Technology: Utilizes IMDG to store data in main memory across multiple computing nodes for distributed in-memory data access.
  • Distributed Processing System: Performs distributed sequence clustering across laboratory clusters and cloud-computing environments such as Amazon EC2.
  • Scalability and Performance: Processes approximately 200 K reads in ~3 hours on a ten-node cluster and one million reads in 11–20 hours on Amazon EC2 using 20–40 nodes.
  • High Accuracy: Demonstrated superior clustering accuracy on 16S rRNA pyrosequences from a mock community compared with DOTUR, mothur, ESPRIT-Tree, UCLUST, and Swarm.

Scientific Applications:

  • Microbial community analysis: Clusters 16S rRNA reads into OTUs to support microbial diversity studies across human, soil, and water microbiomes.

Methodology:

Pre-processing of high-throughput sequencing data followed by distributed clustering of processed 16S rRNA sequences into OTUs using In-Memory Data Grid-enabled computation.

Topics

Details

Tool Type:
desktop application
Operating Systems:
Linux, Windows, Mac
Programming Languages:
Java
Added:
5/5/2018
Last Updated:
12/10/2018

Operations

Publications

Oh J, Choi C, Park M, Kim BK, Hwang K, Lee S, Hong SG, Nasir A, Cho W, Kim KM. CLUSTOM-CLOUD: In-Memory Data Grid-Based Software for Clustering 16S rRNA Sequence Data in the Cloud Environment. PLOS ONE. 2016;11(3):e0151064. doi:10.1371/journal.pone.0151064. PMID:26954507. PMCID:PMC4783016.

Documentation