{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/mimic-it-multi-modal-in-context-instruction","title":"MIMIC-IT: Multi-Modal In-Context Instruction Tuning","arxiv_id":"2306.05425","date":"2023-06-08","proceeding":null,"authors":["Bo Li","Yuanhan Zhang","Liangyu Chen","Jinghao Wang","Fanyi Pu","Jingkang Yang","Chunyuan Li","Ziwei Liu"],"abstract":"High-quality instructions and responses are essential for the zero-shot performance of large language models on interactive natural language tasks. For interactive vision-language tasks involving intricate visual scenes, a large quantity of diverse and creative instruction-response pairs should be imperative to tune vision-language models (VLMs). Nevertheless, the current availability of vision-language instruction-response pairs in terms of quantity, diversity, and creativity remains limited, posing challenges to the generalization of interactive VLMs. Here we present MultI-Modal In-Context Instruction Tuning (MIMIC-IT), a dataset comprising 2.8 million multimodal instruction-response pairs, with 2.2 million unique instructions derived from images and videos. Each pair is accompanied by multi-modal in-context information, forming conversational contexts aimed at empowering VLMs in perception, reasoning, and planning. The instruction-response collection process, dubbed as Syphus, is scaled using an automatic annotation pipeline that combines human expertise with GPT's capabilities. Using the MIMIC-IT dataset, we train a large VLM named Otter. Based on extensive evaluations conducted on vision-language benchmarks, it has been observed that Otter demonstrates remarkable proficiency in multi-modal perception, reasoning, and in-context learning. Human evaluation reveals it effectively aligns with the user's intentions. We release the MIMIC-IT dataset, instruction-response collection pipeline, benchmarks, and the Otter model.","url_abs":"https://arxiv.org/abs/2306.05425v1","url_pdf":"https://arxiv.org/pdf/2306.05425v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"mimic-it-multi-modal-in-context-instruction","repo_url":"https://github.com/luodian/otter","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"mimic-it-multi-modal-in-context-instruction","repo_url":"https://github.com/One-2-3-45/One-2-3-45","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"Apache-2.0"}}],"tasks":[{"task_slug":"in-context-learning","task_name":"In-Context Learning"},{"task_slug":"visual-question-answering-1","task_name":"Visual Question Answering"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/visual-question-answering-on-mm-vet","task":"Visual Question Answering","dataset":"MM-Vet","model":"Otter-9B (MPT-7B)","rank_in_archive_order":222,"of":231,"metrics":{"GPT-4 score":"24.7±0.3","Params":"9B"},"uses_additional_data":false},{"leaderboard":"/sota/visual-question-answering-on-mm-vet","task":"Visual Question Answering","dataset":"MM-Vet","model":"Otter-9B (LLaMA)","rank_in_archive_order":223,"of":231,"metrics":{"GPT-4 score":"24.6±0.2","Params":"9B"},"uses_additional_data":false},{"leaderboard":"/sota/visual-question-answering-on-mm-vet-v2","task":"Visual Question Answering","dataset":"MM-Vet v2","model":"Otter-9B","rank_in_archive_order":23,"of":24,"metrics":{"GPT-4 score":"23.2±0.1","Params":"9B"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/2306.05425","atlas_url":"https://app.syntology.ai/?focus=2306.05425","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}