MedViLL

MedViLL performs multi-modal representation learning to integrate radiological images and their corresponding textual reports for vision-language understanding and generation tasks in radiology.


Key Features:

  • BERT-based Architecture: Extends BERT (Bidirectional Encoder Representations from Transformers) to accept and process multi-modal inputs of radiological images and text.
  • Multi-modal Representation Learning: Learns joint representations that align visual features from radiological images with textual features from reports.
  • Task-tailored Pre-training Objectives: Uses pre-training objectives specifically designed for structured image data and unstructured textual reports.
  • Multi-Modal Attention Masking Scheme: Implements a novel attention masking scheme to manage and align cross-modal interactions between visual and textual tokens.
  • Vision–Language Understanding and Generation: Supports diagnosis classification, medical image-report retrieval, visual question answering, and radiology report generation.

Scientific Applications:

  • Diagnosis Classification: Automates categorization of medical conditions from radiological images using learned multi-modal features.
  • Medical Image-Report Retrieval: Matches radiological images with relevant textual reports via joint image-text representations.
  • Visual Question Answering: Answers clinical questions about radiological images using integrated vision-language context.
  • Radiology Report Generation: Generates radiology reports from images by synthesizing visual features into coherent textual descriptions.

Methodology:

MedViLL extends the BERT architecture with multi-modal representation learning, task-specific pre-training objectives, and a multi-modal attention masking scheme for radiological images and reports, and it was evaluated with statistical comparisons on MIMIC-CXR, Open-I, and VQA-RAD against baseline and task-specific models.

Topics

Details

License:
MIT
Cost:
Free of charge
Tool Type:
command-line tool
Programming Languages:
Python
Added:
11/9/2022
Last Updated:
11/24/2024

Operations

Publications

Moon JH, Lee H, Shin W, Kim Y, Choi E. Multi-Modal Understanding and Generation for Medical Images and Text via Vision-Language Pre-Training. IEEE Journal of Biomedical and Health Informatics. 2022;26(12):6070-6080. doi:10.1109/jbhi.2022.3207502. PMID:36121943.

PMID: 36121943
Funding: - Samsung: IO201211-08109-01 - Institute of Information and Communications Technology Planning and Evaluation: 2019-0-00075 - National Research Foundation of Korea: NRF-2020H1D3A2A03100945