EMBF

EMBF performs de novo short-read clustering to organize erroneous short sequences into a hierarchical tree that models sequencing-error–derived variants and to anchor short gene fragments onto long reference sequences from next-generation sequencing.


Key Features:

  • Frequency-Based Clustering: Clusters short reads by sequence abundance to group sequences likely derived from the same parent despite sequencing errors.
  • Tree Structure Organization: Organizes sequences into a hierarchical tree where each child is considered stochastically derived from a more abundant parent by a single mutation caused by sequencing errors.
  • Anchoring of Short Fragments: Efficiently anchors short gene fragments (30–50 base pairs) onto long reference sequences to support assembly and mapping.
  • Performance: Shows a 3–4× efficiency advantage over SOAP and at least a 150-fold improvement over BLAT, with SOAP exhibiting up to a 30% efficiency degradation on large reference sequences.

Scientific Applications:

  • Genome Assembly: Supports anchoring of short reads to reference sequences to assist genome assembly from next-generation sequencing data.
  • Error Correction: Models sequencing errors as single-mutation derivatives in the tree structure to correct erroneous short reads and improve sequence accuracy.

Methodology:

EMBF applies frequency-based clustering using sequence abundance, constructs a hierarchical tree treating child sequences as stochastic single-mutation derivatives of more abundant parents, and targets short next-generation sequencing reads (30–50 base pairs) for anchoring to long references.

Topics

Details

Tool Type:
web application
Added:
1/13/2017
Last Updated:
11/25/2024

Operations

Data Inputs & Outputs

Publications

Wang W, Zhang P, Liu X. Short read DNA fragment anchoring algorithm. BMC Bioinformatics. 2009;10(S1). doi:10.1186/1471-2105-10-s1-s17. PMID:19208116. PMCID:PMC2648759.