fastseg
fastseg performs rapid segmentation of genomic data to detect copy-number segments and transcript boundaries from DNA microarrays, next-generation sequencing (NGS), and RNA microarray (tiling array) datasets.
Key Features:
- Versatile Segmentation Capabilities: Segments DNA microarrays, NGS data, RNA microarray (tiling arrays) and other matrix or vector-based genomic data to detect copy number segments and transcripts.
- Efficiency and Speed: Provides faster processing compared with earlier methods such as DNACopy while maintaining analytical accuracy.
- Bayesian Coverage Modeling: Uses a Bayesian framework that models depths of coverage at each genomic position with Poisson distributions and mixture components to decompose coverage into integer copy numbers and noise.
- Noise Filtering: Estimates and filters high-noise detections to reduce false positives in segmentation results.
- Benchmark Performance: Demonstrated higher precision (1−FDR) and recall for gains and losses versus five leading CNV methods on simulated datasets, NGS data from a male HapMap individual with implanted CNVs on the X chromosome, known HapMap CNVs, and high-coverage 1000 Genomes Project data.
Scientific Applications:
- Copy Number Variation (CNV) Detection: Quantitative analysis of NGS and microarray data to identify CNVs across genomes.
- Reduction of False Discoveries: Distinguishes true CNVs from noise using Bayesian Poisson mixture modeling and noise filtering to lower FDR.
- Transcript Detection from Tiling Arrays: Identifies transcript boundaries and expression segments in RNA tiling array data via segmentation.
Methodology:
Applies a Bayesian framework modeling per-position depths of coverage with Poisson distributions and mixture components to decompose coverage into integer copy numbers and noise, and includes estimation and filtering of high-noise detections.
Topics
Collections
Details
- License:
- GPL-2.0
- Tool Type:
- command-line tool, library
- Operating Systems:
- Linux, Windows, Mac
- Programming Languages:
- R
- Added:
- 1/17/2017
- Last Updated:
- 12/30/2018
Operations
Data Inputs & Outputs
Sequence analysis
Publications
Klambauer G, Schwarzbauer K, Mayr A, Clevert D, Mitterecker A, Bodenhofer U, Hochreiter S. cn.MOPS: mixture of Poissons for discovering copy number variations in next-generation sequencing data with a low false discovery rate. Nucleic Acids Research. 2012;40(9):e69-e69. doi:10.1093/nar/gks003. PMID:22302147. PMCID:PMC3351174.