BUS-Set

BUS-Set provides a standardized benchmark for evaluating breast ultrasound (BUS) lesion segmentation models using publicly available datasets to improve comparability and reproducibility of model assessments.


Key Features:

  • Publicly available datasets: The benchmark comprises 1,154 BUS images aggregated from four public datasets and includes clinical labels and segmentation annotations.
  • Scanner diversity: Images were acquired using five different scanner types to increase dataset heterogeneity.
  • State-of-the-art architectures: Nine advanced deep learning segmentation architectures were selected for initial benchmarking.
  • Cross-validation: Model performance was assessed using five-fold cross-validation.
  • Statistical evaluation: Model comparisons employed MANOVA/ANOVA with a Tukey post hoc test using a significance threshold of 0.01.
  • Top-performing architecture: Mask R-CNN was identified as the top-performing architecture in the benchmark comparisons.
  • Performance metrics: Reported mean performance metrics include Dice score (0.851), intersection over union (IoU, 0.786), and pixel accuracy (0.975), with a reported comparative p-value > 0.01.
  • Additional dataset evaluation: Mask R-CNN was also evaluated on an additional multi-lesion dataset achieving a mean Dice score of 0.839.
  • Morphological feature analysis: Analysis of regions of interest included Hamming distance, depth-to-width ratio (DWR), circularity, and elongation with reported correlations of 0.888 (DWR), 0.876 (circularity), and 0.532 (elongation) for Mask R-CNN.
  • Training bias considerations: The benchmark examined potential training biases related to lesion size variations within the datasets.
  • Reproducibility: Dataset details and architecture configurations were reported to support reproducible evaluation.

Scientific Applications:

  • Comparative model benchmarking: Enables standardized comparison of BUS lesion segmentation models across multiple architectures and datasets.
  • Diagnostic research and development: Supports development and evaluation of segmentation approaches intended to improve diagnostic accuracy and robustness for breast cancer detection and management.

Methodology:

Nine deep learning segmentation architectures were trained and evaluated with five-fold cross-validation; performance was measured using Dice, IoU, and pixel accuracy; statistical comparisons used MANOVA/ANOVA with Tukey post hoc testing at a 0.01 threshold; morphological analyses computed Hamming distance, DWR, circularity, and elongation with correlation assessment.

Topics

Details

License:
Not licensed
Tool Type:
command-line tool
Programming Languages:
Python
Added:
3/9/2023
Last Updated:
11/24/2024

Operations

Publications

Thomas C, Byra M, Marti R, Yap MH, Zwiggelaar R. BUS‐Set: A benchmark for quantitative evaluation of breast ultrasound segmentation networks with public datasets. Medical Physics. 2023;50(5):3223-3243. doi:10.1002/mp.16287. PMID:36794706.