MoDIL

MoDIL detects medium-sized insertion and deletion polymorphisms (typically 10–50 bp and extending to ~100 bp) from high-throughput clone-end (mate-pair) sequencing data by modeling systematic shifts in mapped-distance distributions to locate and size indels with statistical confidence.


Key Features:

  • Detection range: Targets medium-sized indels typically 10–50 bp and extending to ~100 bp.
  • Input data: Operates on high-throughput clone-end (mate-pair) sequencing data and analyzes clusters of overlapping mate pairs.
  • Reference distribution: Compares the empirical mapped-distance distribution at a locus to the global insert-size distribution p(Y).
  • Mixture modeling: Formulates indel detection as a mixture-of-distributions problem, representing heterozygous events as mixtures of two shifted components and homozygous events as single shifted components.
  • Parameter estimation: Fits component parameters using expectation–maximization with Bayesian regularization.
  • Goodness-of-fit assessment: Evaluates fit using Kolmogorov–Smirnov testing.
  • Error modeling: Uses Gaussian error models with variance that scales inversely with clone coverage to derive indel size estimates and confidence values.
  • Performance: Demonstrates high precision and recall (≥0.85 for ≥20 bp indels) in simulations and whole-genome Illumina datasets.
  • Sensitivity: Identifies variants that are undetectable by callers relying solely on large deviations in insert size.
  • Scalability: Supports genome-wide analysis under high clone coverage.
  • Size estimation accuracy: Predicted indel sizes show strong correlation with validated variation sets and improve as clone depth increases.

Scientific Applications:

  • Medium-sized indel discovery: Detects insertion and deletion polymorphisms in the 10–100 bp range from mate-pair sequencing data.
  • Genotyping: Distinguishes homozygous and heterozygous indels via shifted and mixed distribution components.
  • Whole-genome analysis: Applied to genome-wide Illumina datasets for scalable variant discovery with high clone coverage.
  • Method benchmarking and validation: Evaluated with simulation experiments and comparison to validated variation sets to assess precision, recall, and size-estimation accuracy.

Methodology:

MoDIL models locus-specific mapped-distance distributions as mixtures compared to the global insert-size distribution p(Y), fits components using expectation–maximization with Bayesian regularization, assesses goodness-of-fit with Kolmogorov–Smirnov testing, and derives indel sizes and confidences from Gaussian error models whose variance scales inversely with clone coverage.

Topics

Details

Tool Type:
web application
Operating Systems:
Linux, Windows, Mac
Programming Languages:
Python
Added:
5/2/2017
Last Updated:
11/24/2024

Operations

Publications

Lee S, Hormozdiari F, Alkan C, Brudno M. MoDIL: detecting small indels from clone-end sequencing with mixtures of distributions. Nature Methods. 2009;6(7):473-474. doi:10.1038/nmeth.f.256. PMID:19483690.