fastseg

fastseg performs rapid segmentation of genomic data to detect copy-number segments and transcript boundaries from DNA microarrays, next-generation sequencing (NGS), and RNA microarray (tiling array) datasets.


Key Features:

  • Versatile Segmentation Capabilities: Segments DNA microarrays, NGS data, RNA microarray (tiling arrays) and other matrix or vector-based genomic data to detect copy number segments and transcripts.
  • Efficiency and Speed: Provides faster processing compared with earlier methods such as DNACopy while maintaining analytical accuracy.
  • Bayesian Coverage Modeling: Uses a Bayesian framework that models depths of coverage at each genomic position with Poisson distributions and mixture components to decompose coverage into integer copy numbers and noise.
  • Noise Filtering: Estimates and filters high-noise detections to reduce false positives in segmentation results.
  • Benchmark Performance: Demonstrated higher precision (1−FDR) and recall for gains and losses versus five leading CNV methods on simulated datasets, NGS data from a male HapMap individual with implanted CNVs on the X chromosome, known HapMap CNVs, and high-coverage 1000 Genomes Project data.

Scientific Applications:

  • Copy Number Variation (CNV) Detection: Quantitative analysis of NGS and microarray data to identify CNVs across genomes.
  • Reduction of False Discoveries: Distinguishes true CNVs from noise using Bayesian Poisson mixture modeling and noise filtering to lower FDR.
  • Transcript Detection from Tiling Arrays: Identifies transcript boundaries and expression segments in RNA tiling array data via segmentation.

Methodology:

Applies a Bayesian framework modeling per-position depths of coverage with Poisson distributions and mixture components to decompose coverage into integer copy numbers and noise, and includes estimation and filtering of high-noise detections.

Topics

Collections

Details

License:
GPL-2.0
Tool Type:
command-line tool, library
Operating Systems:
Linux, Windows, Mac
Programming Languages:
R
Added:
1/17/2017
Last Updated:
12/30/2018

Operations

Data Inputs & Outputs

Publications

Klambauer G, Schwarzbauer K, Mayr A, Clevert D, Mitterecker A, Bodenhofer U, Hochreiter S. cn.MOPS: mixture of Poissons for discovering copy number variations in next-generation sequencing data with a low false discovery rate. Nucleic Acids Research. 2012;40(9):e69-e69. doi:10.1093/nar/gks003. PMID:22302147. PMCID:PMC3351174.

Documentation

Downloads

Links