Papers › MIMIC-IT: Multi-Modal In-Context Instruction Tuning

MIMIC-IT: Multi-Modal In-Context Instruction Tuning

8 Jun 2023arXiv:2306.05425archive 2025-07-28

Bo Li, Yuanhan Zhang, Liangyu Chen, Jinghao Wang, Fanyi Pu, Jingkang Yang, Chunyuan Li, Ziwei Liu

High-quality instructions and responses are essential for the zero-shot performance of large language models on interactive natural language tasks. For interactive vision-language tasks involving intricate visual scenes, a large quantity of diverse and creative instruction-response pairs should be imperative to tune vision-language models (VLMs). Nevertheless, the current availability of vision-language instruction-response pairs in terms of quantity, diversity, and creativity remains limited, posing challenges to the generalization of interactive VLMs. Here we present MultI-Modal In-Context Instruction Tuning (MIMIC-IT), a dataset comprising 2.8 million multimodal instruction-response pairs, with 2.2 million unique instructions derived from images and videos. Each pair is accompanied by multi-modal in-context information, forming conversational contexts aimed at empowering VLMs in perception, reasoning, and planning. The instruction-response collection process, dubbed as Syphus, is scaled using an automatic annotation pipeline that combines human expertise with GPT's capabilities. Using the MIMIC-IT dataset, we train a large VLM named Otter. Based on extensive evaluations conducted on vision-language benchmarks, it has been observed that Otter demonstrates remarkable proficiency in multi-modal perception, reasoning, and in-context learning. Human evaluation reveals it effectively aligns with the user's intentions. We release the MIMIC-IT dataset, instruction-response collection pipeline, benchmarks, and the Otter model.

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

luodian/otter officialmentioned in papermentioned on GitHubpytorchMIT report
One-2-3-45/One-2-3-45 mentioned on GitHubpytorchApache-2.0 report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

In-Context LearningVisual Question Answering

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Visual Question Answering MM-Vet Otter-9B (MPT-7B) GPT-4 score 24.7±0.3 #222 of 231 Archive leaderboard report
Visual Question Answering MM-Vet Otter-9B (MPT-7B) Params 9B #222 of 231 Archive leaderboard report
Visual Question Answering MM-Vet Otter-9B (LLaMA) GPT-4 score 24.6±0.2 #223 of 231 Archive leaderboard report
Visual Question Answering MM-Vet Otter-9B (LLaMA) Params 9B #223 of 231 Archive leaderboard report
Visual Question Answering MM-Vet v2 Otter-9B GPT-4 score 23.2±0.1 #23 of 24 Archive leaderboard report
Visual Question Answering MM-Vet v2 Otter-9B Params 9B #23 of 24 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections