Browse State-of-the-Art › Medical Visual Question Answering
Medical Visual Question Answering
47 papers with code · 0 benchmarks · 11 datasets archive 2025-07-28
Benchmarks archive 2025-07-28
No benchmark for this task in the archive.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
11 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 47 papers with code (97 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
30 Jan 2023 17 repositories listed Syntology ran 4 of 8 samples · 4 unverified · 1 pointer-only (licence)The cost of vision-and-language pre-training has become increasingly prohibitive due to end-to-end training of large-scale models.
-
2 Mar 2023 5 repositories listedTherefore, training an effective generalist biomedical model requires high-quality multimodal data, such as parallel image-text pairs.
-
29 Apr 2022 5 repositories listed Syntology ran 18 of 24 samples · 6 unverified · 7 pointer-only (licence)Building models that can be rapidly adapted to novel tasks using only a handful of annotated examples is an open challenge for multimodal machine learning research.
-
7 Mar 2020 5 repositories listed Syntology ran 4 of 4 samples · 0 unverified · 4 pointer-only (licence)To achieve this goal, the first step is to create a visual question answering (VQA) dataset where the AI agent is presented with a pathology image together with a question and is asked to give the correct answer.
-
27 Jul 2023 4 repositories listedHowever, existing models typically have to be fine-tuned on sizeable down-stream datasets, which poses a significant limitation as in many medical applications data is scarce, necessitating models that are capable of…
-
28 Oct 2023 3 repositories listed Syntology ran 12 of 20 samples · 8 unverifiedTo develop our dataset, we first construct two uni-modal resources: 1) The MIMIC-CXR-VQA dataset, our newly created medical visual question answering (VQA) benchmark, specifically designed to augment the imaging…
-
23 Sep 2024 2 repositories listed Syntology ran 12 of 28 samples · 16 unverified · 28 pointer-only (licence)Multimodal Large Language Models (MLLMs) have tremendous potential to improve the accuracy, availability, and cost-effectiveness of healthcare by providing automated solutions or serving as aids to medical professionals.
-
17 May 2023 2 repositories listed Syntology ran 3 of 5 samples · 2 unverifiedMedical Visual Question Answering (MedVQA) presents a significant opportunity to enhance diagnostic accuracy and healthcare delivery by leveraging artificial intelligence to interpret and answer questions based on…
-
13 Mar 2023 2 repositories listed Syntology ran 6 of 13 samples · 7 unverifiedFoundation models trained on large-scale dataset gain a recent surge in CV and NLP.
-
24 Nov 2022 2 repositories listed Syntology ran 6 of 10 samples · 4 unverifiedMedical image visual question answering (VQA) is a task to answer clinical questions, given a radiographic image, which is a challenging problem that requires a model to integrate both vision and language information.
-
19 May 2021 2 repositories listed Syntology ran 4 of 11 samples · 7 unverifiedHowever, most of the existing medical VQA methods rely on external data for transfer learning, while the meta-data within the dataset is not fully utilized.
-
18 Feb 2021 2 repositories listedWe show that SLAKE can be used to facilitate the development and evaluation of Med-VQA systems.
-
26 Sep 2019 2 repositories listedTraditional approaches for Visual Question Answering (VQA) require large amount of labeled data for training.
-
18 May 2025 1 repository listed Syntology ran 1 of 16 samples · 15 unverifiedThe rapid advancement of Large Language Models (LLMs) has stimulated interest in multi-agent collaboration for addressing complex medical tasks.
-
9 Feb 2025 1 repository listedTo address these issues, we introduce the Cross-Modal Clinical Knowledge Distiller (ClinKD), an innovative framework designed to enhance image-text alignment and establish more effective medical knowledge adaptation…
-
18 Dec 2024 1 repository listed Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)Artificial intelligence has advanced in Medical Visual Question Answering (Med-VQA), but prevalent research tends to focus on the accuracy of the answers, often overlooking the reasoning paths and interpretability,…
-
10 Dec 2024 1 repository listedThis paper introduces BiMediX2, a bilingual (Arabic-English) Bio-Medical EXpert Large Multimodal Model (LMM) with a unified architecture that integrates text and visual modalities, enabling advanced image understanding…
-
19 Nov 2024 1 repository listedUnlike general vision-and-language models trained on diverse, non-specialized datasets, MVLMs are purpose-built for the medical domain, automatically extracting and interpreting critical information from medical images…
-
6 Oct 2024 1 repository listed Syntology ran 15 of 15 samples · 0 unverified · 15 pointer-only (licence)In recent advancements, multimodal large language models (MLLMs) have been fine-tuned on specific medical image datasets to address medical visual question answering (Med-VQA) tasks.
-
2 Sep 2024 1 repository listed Syntology ran 0 of 3 samples · 3 unverified · 3 pointer-only (licence)We introduce Kvasir-VQA, an extended dataset derived from the HyperKvasir and Kvasir-Instrument datasets, augmented with question-and-answer annotations to facilitate advanced machine learning tasks in Gastrointestinal…
-
17 Aug 2024 1 repository listed Syntology ran 8 of 8 samples · 0 unverified · 8 pointer-only (licence)This study introduces the Federated Medical Knowledge Injection (FEDMEKI) platform, a new benchmark designed to address the unique challenges of integrating medical knowledge into foundation models under privacy…
-
16 Aug 2024 1 repository listed Syntology ran 8 of 8 samples · 0 unverified · 8 pointer-only (licence)The application of the Multi-modal Large Language Models (MLLMs) in medical clinical scenarios remains underexplored.
-
9 Aug 2024 1 repository listed Syntology ran 4 of 4 samples · 0 unverifiedIn particular, surgical VQA can enhance the interpretation of surgical data, aiding in accurate diagnoses, effective education, and clinical interventions.
-
6 Aug 2024 1 repository listed Syntology ran 8 of 8 samples · 0 unverified · 8 pointer-only (licence)Unlike the existing multimodal datasets, which are limited by the availability of image-text pairs, we have developed the first automated pipeline that scales up multimodal data by generating multigranular visual and…
-
6 Aug 2024 1 repository listedWith growing interest in recent years, medical visual question answering (Med-VQA) has rapidly evolved, with multimodal large language models (MLLMs) emerging as an alternative to classical model architectures.
-
28 Jun 2024 1 repository listed Syntology ran 9 of 19 samples · 10 unverifiedLarge Vision-Language Models (LVLMs) have shown significant potential in assisting medical diagnosis by leveraging extensive biomedical datasets.
-
30 May 2024 1 repository listedThis study reveals that when subjected to simple probing evaluation, state-of-the-art models perform worse than random guessing on medical diagnosis questions.
-
19 Apr 2024 1 repository listed Syntology ran 3 of 7 samples · 4 unverified · 7 pointer-only (licence)In this paper, we propose the Latent Prompt Assist model (LaPA) for medical visual question answering.
-
22 Mar 2024 1 repository listedChest X-ray images are commonly used for predicting acute and chronic cardiopulmonary conditions, but efforts to integrate them with structured clinical data face challenges due to incomplete electronic health records…
-
14 Feb 2024 1 repository listed Syntology ran 7 of 8 samples · 1 unverified · 8 pointer-only (licence)Importantly, all images in this benchmark are sourced from authentic medical scenarios, ensuring alignment with the requirements of the medical field and suitability for evaluating LVLMs.
Syntology lines on 20 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections