RepAHR
RepAHR assembles high-frequency NGS reads to perform de novo repeat identification in eukaryotic genomes, preserving repetitive region structure to improve repeat reconstruction.
Key Features:
- High-Frequency Read Assembly: Assembles high-frequency reads directly rather than relying on high-frequency k-mer assembly to preserve repetitive region structure and improve repeat completeness.
- Efficient Filtering Process: Selects high-frequency reads from NGS datasets using predefined rules derived from high-frequency k-mers to focus assembly on repetitive sequences.
- SPAdes Integration: Uses SPAdes to assemble the filtered high-frequency reads into repeat sequences.
- Performance Evaluation: Evaluated on five datasets against RepARK and REPdenovo using metrics including N50, reference alignment ratio, coverage ratio of the reference genome, and mask ratio of Repbase.
Scientific Applications:
- Structural Variation Detection: Facilitates detection of structural variations by improving identification of repetitive regions.
- Genome Assembly: Improves completeness and accuracy of genome assemblies by reconstructing repeat regions.
Methodology:
Scans NGS reads for high-frequency k-mers, filters reads based on identified high-frequency k-mers using predefined criteria, and assembles the selected high-frequency reads with SPAdes.
Topics
Details
- Programming Languages:
- Java, Python
- Added:
- 1/18/2021
- Last Updated:
- 2/6/2021
Operations
Publications
Liao X, Gao X, Zhang X, Wu F, Wang J. RepAHR: an improved approach for de novo repeat identification by assembly of the high-frequency reads. BMC Bioinformatics. 2020;21(1). doi:10.1186/s12859-020-03779-w. PMID:33076827. PMCID:PMC7574428.
PMID: 33076827
PMCID: PMC7574428
Funding: - National Natural Science Foundation of China: No.61732009, No.61772557
- Hunan Provincial Science and technology Program: No. 2018wk4001
- 111 Project: No.B18059