RepLoc

RepLoc quantifies sequence repetitiveness and detects de novo repetitive elements in genomic sequences using a weighted k-mer coverage approach.


Key Features:

  • Weighted K-Mer Coverage: Implements a weighted k-mer coverage strategy to enhance precision in quantifying sequence repetitiveness and locating repetitive sequences.
  • Repetitiveness Map (Rmap): Generates a repetitiveness map (Rmap) that represents repetitive sequences across the genome for downstream analyses, including segmental duplications.
  • Applications in Genome Analysis: Supports de novo repeat detection, genome structure analysis, and genome mappability analysis based on repetitiveness metrics.
  • De Novo Repeat Detection: Identifies novel repetitive elements within genomic sequences, including common repeats and segmental duplications.
  • Genome Structure and Mappability Analysis: Provides metrics that relate sequence repetitiveness to genome architecture and regions of low mappability.
  • Benchmarking Performance: Demonstrated improved sensitivity and specificity for repeat detection compared with existing k-mer counting-based tools.

Scientific Applications:

  • De Novo Repeat Detection: Detection of novel repetitive elements in assembled genomic sequences.
  • Genome Structure Analysis: Investigation of the relationship between sequence repetitiveness and genome architecture, including segmental duplications.
  • Genome Mappability Analysis: Identification of genomic regions that are challenging to map due to high repetitiveness.
  • Evolution and Regulation Studies: Analysis of repetitive sequences relevant to genome evolution and regulatory processes.

Methodology:

Computes weighted k-mer coverage across input genomic sequences and constructs a repetitiveness map (Rmap); performance was benchmarked against k-mer counting-based tools.

Topics

Details

Added:
1/18/2021
Last Updated:
2/6/2021

Operations

Publications

Feng C, Dai M, Liu Y, Chen M. Sequence repetitiveness quantification and <i>de novo</i> repeat detection by weighted k-mer coverage. Briefings in Bioinformatics. 2020;22(3). doi:10.1093/bib/bbaa086. PMID:32591772.

PMID: 32591772
Funding: - National Key Research and Development Program of China: 2016YFA0501704, 2018YFC0310602 - National Natural Sciences Foundation of China: 31571366, 31771477