CNV-JACG
CNV-JACG assesses and genotypes copy number variations (CNVs) from paired-end whole genome sequencing (WGS) data to evaluate CNV detection accuracy and genotypic attributes.
Key Features:
- Random forest model: Uses an ensemble learning random forest classifier for CNV assessment and genotyping.
- 21 distinctive features: Trains on 21 features that characterize CNV regions and their breakpoints.
- Breakpoint characterization: Incorporates features specific to CNV breakpoints as well as CNV region properties.
- Paired-end WGS input: Operates on paired-end whole genome sequencing data.
- Training datasets: Leverages data from the 1000 Genomes Project, Genome in a Bottle Consortium, Human Genome Structural Variation Consortium, and in-house technical replicates for model training and validation.
- Small CNV sensitivity: Demonstrates superior sensitivity for small CNVs (≤1 kb) compared to alternative genotyping methods.
- Mendelian inconsistency reduction: Reduces Mendelian inconsistencies within trios in comparative analyses.
- Technical-replicate concordance: Improves concordance between technical replicates.
- Comparative benchmarking: Evaluated against the genotyping method SV² for performance comparison.
Scientific Applications:
- CNV genotyping: Assigns genotypes to detected CNVs from WGS data.
- Accuracy assessment: Evaluates the accuracy of CNV calls and reduces false positives/negatives.
- Small-variant discovery: Enhances detection and validation of small CNVs (≤1 kb).
- Trio-based analyses: Identifies and reduces Mendelian inconsistencies in family-based studies.
- Technical reproducibility: Assesses concordance across technical replicates for sequencing experiments.
- Disease-related CNV studies: Supports identification of CNV contributors to human disease and investigations of CNV-linked heritability.
Methodology:
Train a random forest on 21 features describing CNV regions and breakpoints using labeled data from the 1000 Genomes Project, Genome in a Bottle Consortium, Human Genome Structural Variation Consortium, and in-house technical replicates; validate on paired-end WGS by comparing genotyping performance and metrics (Mendelian inconsistencies, technical-replicate concordance) against SV².
Topics
Details
- Tool Type:
- command-line tool
- Programming Languages:
- Perl, R
- Added:
- 3/19/2021
- Last Updated:
- 4/26/2021
Operations
Publications
Zhuang X, Ye R, So M, Lam W, Karim A, Yu M, Ngo ND, Cherny SS, Tam PK, Garcia-Barcelo M, Tang CS, Sham PC. A random forest-based framework for genotyping and accuracy assessment of copy number variations. NAR Genomics and Bioinformatics. 2020;2(3). doi:10.1093/nargab/lqaa071. PMID:33575619. PMCID:PMC7671382.