SICaRiO
SICaRiO enhances detection of true insertions and deletions (indels) in next-generation sequencing (NGS) data to reduce false positives in variant calling.
Key Features:
- Gradient Boosting Classifier: SICaRiO employs a gradient boosting machine learning classifier to classify indel calls.
- Training on Gold-Standard Data: The classifier is trained on the Genome in a Bottle (GIAB) consortium gold-standard dataset.
- Independence from Pipeline-Specific Metrics: The method does not require pipeline-specific information such as read depth and instead uses genomic features derivable from public resources.
- Improved Variant Calling Performance: SICaRiO filters false positive indel calls to improve the performance of variant calling pipelines.
- Genomic Context Insight: The tool analyzes prior genomic contexts that contribute to erroneous indel calls.
- Prediction Difficulty Analysis: SICaRiO compares prediction difficulty across three categories of indels and across sequencing pipelines.
- Feature Ranking for False Positives: The method ranks genomic features by their predictivity for false positive indel calls.
Scientific Applications:
- Personalized Healthcare: Improves accuracy of indel detection for clinical and personalized genomic analyses.
- Disease Genomics: Aids identification of disease-associated indels for research into genetic causes of disease.
- Population Genetics: Enables more accurate studies of genetic diversity and evolutionary patterns by reducing indel call errors.
Methodology:
SICaRiO uses a gradient boosting machine learning classifier trained on GIAB gold-standard data, leverages genomic features derivable from public resources rather than pipeline-specific metrics, compares prediction difficulty across three indel categories and sequencing pipelines, and ranks genomic features by their predictivity for false positives.
Topics
Details
- Added:
- 1/18/2021
- Last Updated:
- 2/18/2021
Operations
Publications
Bhuyan MSI, Pe’er I, Rahman MS. SICaRiO: short indel call filtering with boosting. Briefings in Bioinformatics. 2020;22(4). doi:10.1093/bib/bbaa238. PMID:33003198.
DOI: 10.1093/BIB/BBAA238
PMID: 33003198