Papers › InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
Wenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong, Junqi Zhao, Weisheng Wang, Boyang Li, Pascale Fung, Steven Hoi
Large-scale pre-training and instruction tuning have been successful at creating general-purpose language models with broad competence. However, building general-purpose vision-language models is challenging due to the rich input distributions and task diversity resulting from the additional visual input. Although vision-language pretraining has been widely studied, vision-language instruction tuning remains under-explored. In this paper, we conduct a systematic and comprehensive study on vision-language instruction tuning based on the pretrained BLIP-2 models. We gather 26 publicly available datasets, covering a wide variety of tasks and capabilities, and transform them into instruction tuning format. Additionally, we introduce an instruction-aware Query Transformer, which extracts informative features tailored to the given instruction. Trained on 13 held-in datasets, InstructBLIP attains state-of-the-art zero-shot performance across all 13 held-out datasets, substantially outperforming BLIP-2 and larger Flamingo models. Our models also lead to state-of-the-art performance when finetuned on individual downstream tasks (e.g., 90.7% accuracy on ScienceQA questions with image contexts). Furthermore, we qualitatively demonstrate the advantages of InstructBLIP over concurrent multimodal models. All InstructBLIP models are open-sourced at https://github.com/salesforce/LAVIS/tree/main/projects/instructblip.
In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.
Code
Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| Long-Context Understanding | MMNeedle | InstructBLIP-Flan-T5-XXL | 1 Image, 2*2 Stitching, Exact Accuracy | 3.8 | #8 of 12 | Archive leaderboard | report |
| Long-Context Understanding | MMNeedle | InstructBLIP-Flan-T5-XXL | 1 Image, 4*4 Stitching, Exact Accuracy | 6.2 | #8 of 12 | Archive leaderboard | report |
| Long-Context Understanding | MMNeedle | InstructBLIP-Flan-T5-XXL | 1 Image, 8*8 Stitching, Exact Accuracy | 2.2 | #8 of 12 | Archive leaderboard | report |
| Long-Context Understanding | MMNeedle | InstructBLIP-Flan-T5-XXL | 10 Images, 1*1 Stitching, Exact Accuracy | 0 | #8 of 12 | Archive leaderboard | report |
| Long-Context Understanding | MMNeedle | InstructBLIP-Flan-T5-XXL | 10 Images, 2*2 Stitching, Exact Accuracy | 0 | #8 of 12 | Archive leaderboard | report |
| Long-Context Understanding | MMNeedle | InstructBLIP-Flan-T5-XXL | 10 Images, 4*4 Stitching, Exact Accuracy | 0 | #8 of 12 | Archive leaderboard | report |
| Long-Context Understanding | MMNeedle | InstructBLIP-Flan-T5-XXL | 10 Images, 8*8 Stitching, Exact Accuracy | 0 | #8 of 12 | Archive leaderboard | report |
| Long-Context Understanding | MMNeedle | InstructBLIP-Vicuna-13B | 1 Image, 2*2 Stitching, Exact Accuracy | 0 | #12 of 12 | Archive leaderboard | report |
| Long-Context Understanding | MMNeedle | InstructBLIP-Vicuna-13B | 1 Image, 4*4 Stitching, Exact Accuracy | 0 | #12 of 12 | Archive leaderboard | report |
| Long-Context Understanding | MMNeedle | InstructBLIP-Vicuna-13B | 1 Image, 8*8 Stitching, Exact Accuracy | 0 | #12 of 12 | Archive leaderboard | report |
| Long-Context Understanding | MMNeedle | InstructBLIP-Vicuna-13B | 10 Images, 1*1 Stitching, Exact Accuracy | 0 | #12 of 12 | Archive leaderboard | report |
| Long-Context Understanding | MMNeedle | InstructBLIP-Vicuna-13B | 10 Images, 2*2 Stitching, Exact Accuracy | 0 | #12 of 12 | Archive leaderboard | report |
| Long-Context Understanding | MMNeedle | InstructBLIP-Vicuna-13B | 10 Images, 4*4 Stitching, Exact Accuracy | 0 | #12 of 12 | Archive leaderboard | report |
| Long-Context Understanding | MMNeedle | InstructBLIP-Vicuna-13B | 10 Images, 8*8 Stitching, Exact Accuracy | 0 | #12 of 12 | Archive leaderboard | report |
| Video Question Answering | MVBench | InstructBLIP | Avg. | 32.5 | #21 of 22 | Archive leaderboard | report |
| Visual Question Answering | BenchLMM | InstructBLIP-13B | GPT-3.5 score | 45.03 | #5 of 10 | Archive leaderboard | report |
| Visual Question Answering | BenchLMM | InstructBLIP-7B | GPT-3.5 score | 44.63 | #6 of 10 | Archive leaderboard | report |
| Visual Question Answering | ViP-Bench | InstructBLIP-13B (Visual Prompt) | GPT-4 score (bbox) | 35.8 | #10 of 13 | Archive leaderboard | report |
| Visual Question Answering | ViP-Bench | InstructBLIP-13B (Visual Prompt) | GPT-4 score (human) | 35.2 | #10 of 13 | Archive leaderboard | report |
| Visual Question Answering (VQA) | InfiMM-Eval | InstructBLIP | Abductive | 37.76 | #8 of 14 | Archive leaderboard | report |
| Visual Question Answering (VQA) | InfiMM-Eval | InstructBLIP | Analogical | 20.56 | #8 of 14 | Archive leaderboard | report |
| Visual Question Answering (VQA) | InfiMM-Eval | InstructBLIP | Deductive | 27.56 | #8 of 14 | Archive leaderboard | report |
| Visual Question Answering (VQA) | InfiMM-Eval | InstructBLIP | Overall score | 28.02 | #8 of 14 | Archive leaderboard | report |
| Visual Question Answering (VQA) | InfiMM-Eval | InstructBLIP | Params | 8B | #8 of 14 | Archive leaderboard | report |
| visual instruction following | LLaVA-Bench | InstructBLIP-7B | avg score | 60.9 | #6 of 8 | Archive leaderboard | report |
| visual instruction following | LLaVA-Bench | InstructBLIP-13B | avg score | 58.2 | #7 of 8 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Methods
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections