CRISP
CRISP identifies rare and common single nucleotide polymorphisms (SNPs) from pooled next-generation sequencing datasets to enable accurate variant discovery across large populations.
Key Features:
- Variant Identification: CRISP detects rare and common single nucleotide polymorphisms (SNPs) from pooled sequencing data.
- Statistical Approach: It compares allele counts across multiple pools using contingency tables and evaluates the probability of observing non-reference base calls to distinguish true variants from sequencing errors.
- Data Integration: The method incorporates distribution of reads between forward and reverse strands and uses pool size information to filter false positives.
- Validation and Performance: Validation on two pooled sequencing datasets generated using the Illumina Genome Analyzer demonstrated detection of 80-85% of SNPs identified by individual sequencing and a false discovery rate of 3-5%.
- Comparative Advantage: Compared with previous pooled SNP detection methods, CRISP exhibits significantly lower false positive and false negative rates.
Scientific Applications:
- Population-scale variant screening: Enables cost-effective analysis of large populations using pooled sequencing.
- Rare variant discovery: Facilitates identification of rare genetic variants for studies of genetic diversity and population genetics.
- Disease association studies: Supports pooled sequencing–based association studies to detect variant-disease correlations.
Methodology:
Computational methods include contingency table analysis to compare allele frequencies across pools; assessment of the probability of sequencing errors underlying observed non-reference base calls; and use of strand-distribution and pool-size information to filter false positives.
Topics
Details
- Tool Type:
- command-line tool
- Programming Languages:
- Python, C
- Added:
- 1/13/2017
- Last Updated:
- 11/25/2024
Operations
Data Inputs & Outputs
Genetic variation analysis
Inputs
Outputs
Publications
Bansal V. A statistical method for the detection of variants from next-generation resequencing of DNA pools. Bioinformatics. 2010;26(12):i318-i324. doi:10.1093/bioinformatics/btq214. PMID:20529923. PMCID:PMC2881398.