ISSAKE

ISSAKE assembles T-cell receptor (TCR) beta-chain CDR3 sequences from short-read massively parallel sequencing of peripheral blood to profile TCR diversity at sequence-level resolution.


Key Features:

  • Sequencing approach: Tailored for massively parallel sequencing of T-cell metagenomes using short reads and targets the CDR3 region of the TCR beta-chain.
  • De novo assembly: Employs de novo assembly that reconstructs full-length sequences from short reads, using reads that align to known V genes with consecutive unmatched bases in the adjacent CDR3 region as assembly seeds.
  • K-mer search and 3' read extension: Leverages k-mer search techniques combined with 3' read extension to assemble CDR3-containing sequences from short reads.
  • Error modeling and simulation: Incorporates models of sequence diversity and error distribution to simulate distinct TCR clonotypes at frequencies of 1–20 parts per million and to generate reads with controlled random errors (1–2%).
  • Sensitivity and detection: Reports sensitivity metrics for assembled read lengths: 36 nt reads detect >51% of 1 ppm clonotypes with random errors and 63% with modeled errors; 42 nt and 50 nt assembled reads detect 82.0% and 94.7% of 1 ppm clonotypes respectively, and clonotypes at ≥6 ppm are detected at >99%.

Scientific Applications:

  • TCR repertoire profiling: Enables high-resolution, sequence-level profiling of TCR repertoires from peripheral blood.
  • Vaccine and immunotherapy research: Supports identification and quantification of clonotypes relevant to vaccine responses and immunotherapeutic interventions.
  • Disease pathogenesis and immune response studies: Facilitates analysis of T-cell diversity and clonotype dynamics in studies of disease pathogenesis and immune system function.

Methodology:

De novo assembly using k-mer search and 3' read extension; use of reads aligning to known V genes with consecutive unmatched CDR3 bases as seeds; modeling of sequence diversity and error distributions; simulation of clonotypes at 1–20 ppm with reads containing 1–2% random errors; evaluation of detection sensitivity for assembled read lengths of 36, 42, and 50 nt.

Topics

Details

License:
GPL-2.0
Maturity:
Legacy
Tool Type:
command-line tool
Programming Languages:
Perl, Python
Added:
1/13/2017
Last Updated:
11/25/2024

Operations

Publications

Warren RL, Nelson BH, Holt RA. Profiling model T-cell metagenomes with short reads. Bioinformatics. 2009;25(4):458-464. doi:10.1093/bioinformatics/btp010. PMID:19136549.