CRISP

CRISP identifies rare and common single nucleotide polymorphisms (SNPs) from pooled next-generation sequencing datasets to enable accurate variant discovery across large populations.


Key Features:

  • Variant Identification: CRISP detects rare and common single nucleotide polymorphisms (SNPs) from pooled sequencing data.
  • Statistical Approach: It compares allele counts across multiple pools using contingency tables and evaluates the probability of observing non-reference base calls to distinguish true variants from sequencing errors.
  • Data Integration: The method incorporates distribution of reads between forward and reverse strands and uses pool size information to filter false positives.
  • Validation and Performance: Validation on two pooled sequencing datasets generated using the Illumina Genome Analyzer demonstrated detection of 80-85% of SNPs identified by individual sequencing and a false discovery rate of 3-5%.
  • Comparative Advantage: Compared with previous pooled SNP detection methods, CRISP exhibits significantly lower false positive and false negative rates.

Scientific Applications:

  • Population-scale variant screening: Enables cost-effective analysis of large populations using pooled sequencing.
  • Rare variant discovery: Facilitates identification of rare genetic variants for studies of genetic diversity and population genetics.
  • Disease association studies: Supports pooled sequencing–based association studies to detect variant-disease correlations.

Methodology:

Computational methods include contingency table analysis to compare allele frequencies across pools; assessment of the probability of sequencing errors underlying observed non-reference base calls; and use of strand-distribution and pool-size information to filter false positives.

Topics

Details

Tool Type:
command-line tool
Programming Languages:
Python, C
Added:
1/13/2017
Last Updated:
11/25/2024

Operations

Data Inputs & Outputs

Genetic variation analysis

Publications

Bansal V. A statistical method for the detection of variants from next-generation resequencing of DNA pools. Bioinformatics. 2010;26(12):i318-i324. doi:10.1093/bioinformatics/btq214. PMID:20529923. PMCID:PMC2881398.

Documentation