RepAHR

RepAHR assembles high-frequency NGS reads to perform de novo repeat identification in eukaryotic genomes, preserving repetitive region structure to improve repeat reconstruction.


Key Features:

  • High-Frequency Read Assembly: Assembles high-frequency reads directly rather than relying on high-frequency k-mer assembly to preserve repetitive region structure and improve repeat completeness.
  • Efficient Filtering Process: Selects high-frequency reads from NGS datasets using predefined rules derived from high-frequency k-mers to focus assembly on repetitive sequences.
  • SPAdes Integration: Uses SPAdes to assemble the filtered high-frequency reads into repeat sequences.
  • Performance Evaluation: Evaluated on five datasets against RepARK and REPdenovo using metrics including N50, reference alignment ratio, coverage ratio of the reference genome, and mask ratio of Repbase.

Scientific Applications:

  • Structural Variation Detection: Facilitates detection of structural variations by improving identification of repetitive regions.
  • Genome Assembly: Improves completeness and accuracy of genome assemblies by reconstructing repeat regions.

Methodology:

Scans NGS reads for high-frequency k-mers, filters reads based on identified high-frequency k-mers using predefined criteria, and assembles the selected high-frequency reads with SPAdes.

Topics

Details

Programming Languages:
Java, Python
Added:
1/18/2021
Last Updated:
2/6/2021

Operations

Publications

Liao X, Gao X, Zhang X, Wu F, Wang J. RepAHR: an improved approach for de novo repeat identification by assembly of the high-frequency reads. BMC Bioinformatics. 2020;21(1). doi:10.1186/s12859-020-03779-w. PMID:33076827. PMCID:PMC7574428.

PMID: 33076827
PMCID: PMC7574428
Funding: - National Natural Science Foundation of China: No.61732009, No.61772557 - Hunan Provincial Science and technology Program: No. 2018wk4001 - 111 Project: No.B18059