Papers › InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

11 May 2023NeurIPS 2023 11arXiv:2305.06500archive 2025-07-28

Wenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong, Junqi Zhao, Weisheng Wang, Boyang Li, Pascale Fung, Steven Hoi

Large-scale pre-training and instruction tuning have been successful at creating general-purpose language models with broad competence. However, building general-purpose vision-language models is challenging due to the rich input distributions and task diversity resulting from the additional visual input. Although vision-language pretraining has been widely studied, vision-language instruction tuning remains under-explored. In this paper, we conduct a systematic and comprehensive study on vision-language instruction tuning based on the pretrained BLIP-2 models. We gather 26 publicly available datasets, covering a wide variety of tasks and capabilities, and transform them into instruction tuning format. Additionally, we introduce an instruction-aware Query Transformer, which extracts informative features tailored to the given instruction. Trained on 13 held-in datasets, InstructBLIP attains state-of-the-art zero-shot performance across all 13 held-out datasets, substantially outperforming BLIP-2 and larger Flamingo models. Our models also lead to state-of-the-art performance when finetuned on individual downstream tasks (e.g., 90.7% accuracy on ScienceQA questions with image contexts). Furthermore, we qualitatively demonstrate the advantages of InstructBLIP over concurrent multimodal models. All InstructBLIP models are open-sourced at https://github.com/salesforce/LAVIS/tree/main/projects/instructblip.

PaperPDFConference PDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

salesforce/lavis officialmentioned in papermentioned on GitHubpytorchBSD-3-Clause report
tabtoyou/kollava mentioned on GitHubpytorch report
MS-P3/code3 mindsporeApache-2.0 report
pwc-1/Paper-9 mindspore report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

1 Image, 2*2 StitchingDiversityImage RetrievalLong-Context UnderstandingVideo Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)visual instruction following

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Long-Context Understanding MMNeedle InstructBLIP-Flan-T5-XXL 1 Image, 2*2 Stitching, Exact Accuracy 3.8 #8 of 12 Archive leaderboard report
Long-Context Understanding MMNeedle InstructBLIP-Flan-T5-XXL 1 Image, 4*4 Stitching, Exact Accuracy 6.2 #8 of 12 Archive leaderboard report
Long-Context Understanding MMNeedle InstructBLIP-Flan-T5-XXL 1 Image, 8*8 Stitching, Exact Accuracy 2.2 #8 of 12 Archive leaderboard report
Long-Context Understanding MMNeedle InstructBLIP-Flan-T5-XXL 10 Images, 1*1 Stitching, Exact Accuracy 0 #8 of 12 Archive leaderboard report
Long-Context Understanding MMNeedle InstructBLIP-Flan-T5-XXL 10 Images, 2*2 Stitching, Exact Accuracy 0 #8 of 12 Archive leaderboard report
Long-Context Understanding MMNeedle InstructBLIP-Flan-T5-XXL 10 Images, 4*4 Stitching, Exact Accuracy 0 #8 of 12 Archive leaderboard report
Long-Context Understanding MMNeedle InstructBLIP-Flan-T5-XXL 10 Images, 8*8 Stitching, Exact Accuracy 0 #8 of 12 Archive leaderboard report
Long-Context Understanding MMNeedle InstructBLIP-Vicuna-13B 1 Image, 2*2 Stitching, Exact Accuracy 0 #12 of 12 Archive leaderboard report
Long-Context Understanding MMNeedle InstructBLIP-Vicuna-13B 1 Image, 4*4 Stitching, Exact Accuracy 0 #12 of 12 Archive leaderboard report
Long-Context Understanding MMNeedle InstructBLIP-Vicuna-13B 1 Image, 8*8 Stitching, Exact Accuracy 0 #12 of 12 Archive leaderboard report
Long-Context Understanding MMNeedle InstructBLIP-Vicuna-13B 10 Images, 1*1 Stitching, Exact Accuracy 0 #12 of 12 Archive leaderboard report
Long-Context Understanding MMNeedle InstructBLIP-Vicuna-13B 10 Images, 2*2 Stitching, Exact Accuracy 0 #12 of 12 Archive leaderboard report
Long-Context Understanding MMNeedle InstructBLIP-Vicuna-13B 10 Images, 4*4 Stitching, Exact Accuracy 0 #12 of 12 Archive leaderboard report
Long-Context Understanding MMNeedle InstructBLIP-Vicuna-13B 10 Images, 8*8 Stitching, Exact Accuracy 0 #12 of 12 Archive leaderboard report
Video Question Answering MVBench InstructBLIP Avg. 32.5 #21 of 22 Archive leaderboard report
Visual Question Answering BenchLMM InstructBLIP-13B GPT-3.5 score 45.03 #5 of 10 Archive leaderboard report
Visual Question Answering BenchLMM InstructBLIP-7B GPT-3.5 score 44.63 #6 of 10 Archive leaderboard report
Visual Question Answering ViP-Bench InstructBLIP-13B (Visual Prompt) GPT-4 score (bbox) 35.8 #10 of 13 Archive leaderboard report
Visual Question Answering ViP-Bench InstructBLIP-13B (Visual Prompt) GPT-4 score (human) 35.2 #10 of 13 Archive leaderboard report
Visual Question Answering (VQA) InfiMM-Eval InstructBLIP Abductive 37.76 #8 of 14 Archive leaderboard report
Visual Question Answering (VQA) InfiMM-Eval InstructBLIP Analogical 20.56 #8 of 14 Archive leaderboard report
Visual Question Answering (VQA) InfiMM-Eval InstructBLIP Deductive 27.56 #8 of 14 Archive leaderboard report
Visual Question Answering (VQA) InfiMM-Eval InstructBLIP Overall score 28.02 #8 of 14 Archive leaderboard report
Visual Question Answering (VQA) InfiMM-Eval InstructBLIP Params 8B #8 of 14 Archive leaderboard report
visual instruction following LLaVA-Bench InstructBLIP-7B avg score 60.9 #6 of 8 Archive leaderboard report
visual instruction following LLaVA-Bench InstructBLIP-13B avg score 58.2 #7 of 8 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

Absolute Position EncodingsAdamAttentionBPEDense ConnectionsDropoutLabel SmoothingLayer NormalizationLinear LayerMulti-Head AttentionPosition-Wise Feed-Forward LayerResidual ConnectionSoftmaxTransformer

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections