DeepSVFilter

DeepSVFilter filters false positive structural variant (SV) calls from short read whole genome sequencing using deep learning to improve SV call specificity.


Key Features:

  • Deep Learning Framework: Encodes structural variant signals from read alignments into image formats and classifies them with convolutional neural networks (CNNs).
  • Transfer Learning with Pre-trained CNNs: Employs transfer learning with pre-trained CNN models trained on well-characterized samples containing high-confidence structural variants to improve generalization to new datasets.
  • Versatility in Application: Operates as a standalone SV filter or integrates with existing structural variant detection tools to enhance specificity of SV calls.
  • Performance Validation: Validated on two well-characterized samples, demonstrating reduction of false positive SV calls when coupled with existing SV detection methods.
  • Implementation: Implemented in Python.

Scientific Applications:

  • Genetic disease studies: Enhances accuracy of SV calls used to identify disease-associated structural variants in research cohorts.
  • Cancer genomics: Reduces false positive SVs in tumor genome analyses to improve detection of somatic structural alterations.
  • Population genetics: Improves specificity of SV catalogs derived from short read whole genome sequencing for population-scale studies.
  • Clinical genomics: Supports interpretation of SVs in clinical sequencing by lowering false positive rates.

Methodology:

Encodes SV signals from read alignments into image formats and classifies them with convolutional neural networks (CNNs), using transfer learning with pre-trained CNN models trained on well-characterized samples containing high-confidence structural variants.

Topics

Details

License:
MIT
Tool Type:
command-line tool
Programming Languages:
Python
Added:
1/18/2021
Last Updated:
2/27/2021

Operations

Publications

Liu Y, Huang Y, Wang G, Wang Y. A deep learning approach for filtering structural variants in short read sequencing data. Briefings in Bioinformatics. 2020;22(4). doi:10.1093/bib/bbaa370. PMID:33378767.

PMID: 33378767
Funding: - National Key R&D Program of China: 2017YFC0907500 - Fundamental Research Funds for the Central Universities: HIT.NSRIF.2019055 - Heilongjiang Postdoctoral Financial Assistance: LBH-Z17070 - China Postdoctoral Science Foundation: 2018 M631934, 2018 T110300 - Natural Science Foundation of China: 31701147