Papers › LLaVA-Med: Training a Large Language-and-Vision Assistant for Biomedicine in One Day

LLaVA-Med: Training a Large Language-and-Vision Assistant for Biomedicine in One Day

1 Jun 2023NeurIPS 2023 11arXiv:2306.00890archive 2025-07-28

Chunyuan Li, Cliff Wong, Sheng Zhang, Naoto Usuyama, Haotian Liu, Jianwei Yang, Tristan Naumann, Hoifung Poon, Jianfeng Gao

Conversational generative AI has demonstrated remarkable promise for empowering biomedical practitioners, but current investigations focus on unimodal text. Multimodal conversational AI has seen rapid progress by leveraging billions of image-text pairs from the public web, but such general-domain vision-language models still lack sophistication in understanding and conversing about biomedical images. In this paper, we propose a cost-efficient approach for training a vision-language conversational assistant that can answer open-ended research questions of biomedical images. The key idea is to leverage a large-scale, broad-coverage biomedical figure-caption dataset extracted from PubMed Central, use GPT-4 to self-instruct open-ended instruction-following data from the captions, and then fine-tune a large general-domain vision-language model using a novel curriculum learning method. Specifically, the model first learns to align biomedical vocabulary using the figure-caption pairs as is, then learns to master open-ended conversational semantics using GPT-4 generated instruction-following data, broadly mimicking how a layperson gradually acquires biomedical knowledge. This enables us to train a Large Language and Vision Assistant for BioMedicine (LLaVA-Med) in less than 15 hours (with eight A100s). LLaVA-Med exhibits excellent multimodal conversational capability and can follow open-ended instruction to assist with inquiries about a biomedical image. On three standard biomedical visual question answering datasets, LLaVA-Med outperforms previous supervised state-of-the-art on certain metrics. To facilitate biomedical multimodal research, we will release our instruction-following data and the LLaVA-Med model.

PaperPDFConference PDFCode

In Syntology View this paper on Syntology: its repositories, every harvested function with whether it ran, its licence and the call to fetch it.

Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

microsoft/LLaVA-Med mentioned on GitHubpytorchNOASSERTION report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Image ClassificationInstruction FollowingLanguage ModellingQuestion AnsweringReferring Expression ComprehensionReferring expression generationVisual Question Answering

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Image Classification ColonINST-v1 (Seen) LLaVA-Med-v1.0 (w/o LoRA, w/ extra data) Accuray 93.84 #2 of 17 Archive leaderboard report
Image Classification ColonINST-v1 (Seen) LLaVA-Med-v1.5 (w/ LoRA, w/o extra data) Accuray 93.62 #4 of 17 Archive leaderboard report
Image Classification ColonINST-v1 (Seen) LLaVA-Med-v1.0 (w/o LoRA, w/o extra data) Accuray 93.52 #5 of 17 Archive leaderboard report
Image Classification ColonINST-v1 (Seen) LLaVA-Med-v1.5 (w/ LoRA, w/ extra data) Accuray 87.22 #17 of 17 Archive leaderboard report
Image Classification ColonINST-v1 (Unseen) LLaVA-Med-v1.5 (w/ LoRA, w/o extra data) Accuray 79.24 #5 of 17 Archive leaderboard report
Image Classification ColonINST-v1 (Unseen) LLaVA-Med-v1.0 (w/o LoRA, w/o extra data) Accuray 78.04 #10 of 17 Archive leaderboard report
Image Classification ColonINST-v1 (Unseen) LLaVA-Med-v1.0 (w/o LoRA, w/ extra data) Accuray 77.38 #12 of 17 Archive leaderboard report
Image Classification ColonINST-v1 (Unseen) LLaVA-Med-v1.5 (w/ LoRA, w/ extra data) Accuray 66.51 #16 of 17 Archive leaderboard report
Referring expression generation ColonINST-v1 (Seen) LLaVA-Med-v1.5 (w/ LoRA, w/o extra data) Accuray 99.3 #3 of 17 Archive leaderboard report
Referring expression generation ColonINST-v1 (Seen) LLaVA-Med-v1.0 (w/o LoRA, w/o extra data) Accuray 97.74 #9 of 17 Archive leaderboard report
Referring expression generation ColonINST-v1 (Seen) LLaVA-Med-v1.0 (w/o LoRA, w/ extra data) Accuray 97.35 #10 of 17 Archive leaderboard report
Referring expression generation ColonINST-v1 (Seen) LLaVA-Med-v1.5 (w/ LoRA, w/ extra data) Accuray 90.4 #14 of 17 Archive leaderboard report
Referring expression generation ColonINST-v1 (Unseen) LLaVA-Med-v1.0 (w/o LoRA, w/ extra data) Accuray 75.25 #3 of 17 Archive leaderboard report
Referring expression generation ColonINST-v1 (Unseen) LLaVA-Med-v1.0 (w/o LoRA, w/o extra data) Accuray 75.07 #5 of 17 Archive leaderboard report
Referring expression generation ColonINST-v1 (Unseen) LLaVA-Med-v1.5 (w/ LoRA, w/o extra data) Accuray 73.05 #8 of 17 Archive leaderboard report
Referring expression generation ColonINST-v1 (Unseen) LLaVA-Med-v1.5 (w/ LoRA, w/ extra data) Accuray 70.00 #13 of 17 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

ALIGNAbsolute Position EncodingsAdamAttentionBPEDense ConnectionsDropoutFocusGPT-4Label SmoothingLayer NormalizationLinear LayerMulti-Head AttentionPosition-Wise Feed-Forward LayerResidual ConnectionSoftmaxTransformer

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections