Papers › Otter: A Multi-Modal Model with In-Context Instruction Tuning

Otter: A Multi-Modal Model with In-Context Instruction Tuning

5 May 2023arXiv:2305.03726archive 2025-07-28

Bo Li, Yuanhan Zhang, Liangyu Chen, Jinghao Wang, Jingkang Yang, Ziwei Liu

Large language models (LLMs) have demonstrated significant universal capabilities as few/zero-shot learners in various tasks due to their pre-training on vast amounts of text data, as exemplified by GPT-3, which boosted to InstrctGPT and ChatGPT, effectively following natural language instructions to accomplish real-world tasks. In this paper, we propose to introduce instruction tuning into multi-modal models, motivated by the Flamingo model's upstream interleaved format pretraining dataset. We adopt a similar approach to construct our MultI-Modal In-Context Instruction Tuning (MIMIC-IT) dataset. We then introduce Otter, a multi-modal model based on OpenFlamingo (open-sourced version of DeepMind's Flamingo), trained on MIMIC-IT and showcasing improved instruction-following ability and in-context learning. We also optimize OpenFlamingo's implementation for researchers, democratizing the required training resources from 1× A100 GPU to 4× RTX-3090 GPUs, and integrate both OpenFlamingo and Otter into Huggingface Transformers for more researchers to incorporate the models into their customized training and inference pipelines.

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

luodian/otter officialmentioned in papermentioned on GitHubpytorchMIT report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

In-Context LearningInstruction FollowingVisual Question AnsweringVisual Question Answering (VQA)Visual Reasoning

1 archive task tag without a task page not shown.

Datasets

Introduced by this paper, per the archive.

MIMIC-IT

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Visual Question Answering BenchLMM Otter-7B GPT-3.5 score 39.13 #8 of 10 Archive leaderboard report
Visual Question Answering (VQA) InfiMM-Eval Otter Abductive 33.64 #10 of 14 Archive leaderboard report
Visual Question Answering (VQA) InfiMM-Eval Otter Analogical 13.33 #10 of 14 Archive leaderboard report
Visual Question Answering (VQA) InfiMM-Eval Otter Deductive 22.49 #10 of 14 Archive leaderboard report
Visual Question Answering (VQA) InfiMM-Eval Otter Overall score 22.69 #10 of 14 Archive leaderboard report
Visual Question Answering (VQA) InfiMM-Eval Otter Params 7B #10 of 14 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

AdamAttentionAttention DropoutBPECosine AnnealingDense ConnectionsDropoutGPT-3Layer NormalizationLinear LayerLinear Warmup With Cosine AnnealingMulti-Head AttentionResidual ConnectionSoftmaxWeight Decay

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections