MoDIL
MoDIL detects medium-sized insertion and deletion polymorphisms (typically 10–50 bp and extending to ~100 bp) from high-throughput clone-end (mate-pair) sequencing data by modeling systematic shifts in mapped-distance distributions to locate and size indels with statistical confidence.
Key Features:
- Detection range: Targets medium-sized indels typically 10–50 bp and extending to ~100 bp.
- Input data: Operates on high-throughput clone-end (mate-pair) sequencing data and analyzes clusters of overlapping mate pairs.
- Reference distribution: Compares the empirical mapped-distance distribution at a locus to the global insert-size distribution p(Y).
- Mixture modeling: Formulates indel detection as a mixture-of-distributions problem, representing heterozygous events as mixtures of two shifted components and homozygous events as single shifted components.
- Parameter estimation: Fits component parameters using expectation–maximization with Bayesian regularization.
- Goodness-of-fit assessment: Evaluates fit using Kolmogorov–Smirnov testing.
- Error modeling: Uses Gaussian error models with variance that scales inversely with clone coverage to derive indel size estimates and confidence values.
- Performance: Demonstrates high precision and recall (≥0.85 for ≥20 bp indels) in simulations and whole-genome Illumina datasets.
- Sensitivity: Identifies variants that are undetectable by callers relying solely on large deviations in insert size.
- Scalability: Supports genome-wide analysis under high clone coverage.
- Size estimation accuracy: Predicted indel sizes show strong correlation with validated variation sets and improve as clone depth increases.
Scientific Applications:
- Medium-sized indel discovery: Detects insertion and deletion polymorphisms in the 10–100 bp range from mate-pair sequencing data.
- Genotyping: Distinguishes homozygous and heterozygous indels via shifted and mixed distribution components.
- Whole-genome analysis: Applied to genome-wide Illumina datasets for scalable variant discovery with high clone coverage.
- Method benchmarking and validation: Evaluated with simulation experiments and comparison to validated variation sets to assess precision, recall, and size-estimation accuracy.
Methodology:
MoDIL models locus-specific mapped-distance distributions as mixtures compared to the global insert-size distribution p(Y), fits components using expectation–maximization with Bayesian regularization, assesses goodness-of-fit with Kolmogorov–Smirnov testing, and derives indel sizes and confidences from Gaussian error models whose variance scales inversely with clone coverage.
Topics
Details
- Tool Type:
- web application
- Operating Systems:
- Linux, Windows, Mac
- Programming Languages:
- Python
- Added:
- 5/2/2017
- Last Updated:
- 11/24/2024
Operations
Publications
Lee S, Hormozdiari F, Alkan C, Brudno M. MoDIL: detecting small indels from clone-end sequencing with mixtures of distributions. Nature Methods. 2009;6(7):473-474. doi:10.1038/nmeth.f.256. PMID:19483690.