CNAnorm

CNAnorm estimates copy number alterations (CNA) from low-coverage high-throughput sequencing data (one read every 100–10,000 base pairs), correcting for normal-cell contamination and genome size variation to enable ploidy-aware CNA calling.


Key Features:

  • Low Coverage Sequencing Adaptation: Tailored to low coverage data (one read every 100–10,000 base pairs) for CNA analysis from sparse sequencing.
  • Ratio and GC Content Correction: Performs corrections for ratio biases and GC content variations to improve CNA accuracy.
  • Normalization Methodology: Employs a "discrete" normalization approach that analyzes the multi-modality of smoothed ratio distributions to identify modes reflecting ploidy states and contamination levels.
  • Tumor Content Estimation: Identifies at least two distinct ploidy states to provide insights into tumor cell fraction within samples.
  • Correction for Normal Cell Contamination: Corrects for the presence of normal cells in tumor samples to recover true CNA signals.
  • Genome Size Adjustment: Adjusts for differences in genome size to ensure accurate copy number estimates across regions and sample types.

Scientific Applications:

  • Cancer Genomics: Distinguishes between normal and cancerous cell populations within tumor samples to support analyses of tumor heterogeneity and evolution.
  • Public Data Analysis: Applied to publicly available cell line datasets such as HCC1143 and COLO829 for benchmarking and broader genomic studies.

Methodology:

Identifies the multi-modal distribution of smoothed ratios, estimates underlying ploidy levels and normal-cell contamination from the identified modes, and applies corrections to adjust CNA estimates accordingly.

Topics

Collections

Details

License:
GPL-2.0
Tool Type:
command-line tool, library
Operating Systems:
Linux, Windows, Mac
Programming Languages:
R
Added:
1/17/2017
Last Updated:
12/30/2018

Operations

Publications

Gusnanto A, Wood HM, Pawitan Y, Rabbitts P, Berri S. Correcting for cancer genome size and tumour cell content enables better estimation of copy number alterations from next-generation sequence data. Bioinformatics. 2011;28(1):40-47. doi:10.1093/bioinformatics/btr593. PMID:22039209.

Documentation

Downloads

Links