MedViLL
MedViLL performs multi-modal representation learning to integrate radiological images and their corresponding textual reports for vision-language understanding and generation tasks in radiology.
Key Features:
- BERT-based Architecture: Extends BERT (Bidirectional Encoder Representations from Transformers) to accept and process multi-modal inputs of radiological images and text.
- Multi-modal Representation Learning: Learns joint representations that align visual features from radiological images with textual features from reports.
- Task-tailored Pre-training Objectives: Uses pre-training objectives specifically designed for structured image data and unstructured textual reports.
- Multi-Modal Attention Masking Scheme: Implements a novel attention masking scheme to manage and align cross-modal interactions between visual and textual tokens.
- Vision–Language Understanding and Generation: Supports diagnosis classification, medical image-report retrieval, visual question answering, and radiology report generation.
Scientific Applications:
- Diagnosis Classification: Automates categorization of medical conditions from radiological images using learned multi-modal features.
- Medical Image-Report Retrieval: Matches radiological images with relevant textual reports via joint image-text representations.
- Visual Question Answering: Answers clinical questions about radiological images using integrated vision-language context.
- Radiology Report Generation: Generates radiology reports from images by synthesizing visual features into coherent textual descriptions.
Methodology:
MedViLL extends the BERT architecture with multi-modal representation learning, task-specific pre-training objectives, and a multi-modal attention masking scheme for radiological images and reports, and it was evaluated with statistical comparisons on MIMIC-CXR, Open-I, and VQA-RAD against baseline and task-specific models.
Topics
Details
- License:
- MIT
- Cost:
- Free of charge
- Tool Type:
- command-line tool
- Programming Languages:
- Python
- Added:
- 11/9/2022
- Last Updated:
- 11/24/2024
Operations
Publications
Moon JH, Lee H, Shin W, Kim Y, Choi E. Multi-Modal Understanding and Generation for Medical Images and Text via Vision-Language Pre-Training. IEEE Journal of Biomedical and Health Informatics. 2022;26(12):6070-6080. doi:10.1109/jbhi.2022.3207502. PMID:36121943.