{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/unified-io-2-scaling-autoregressive","title":"Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action","arxiv_id":"2312.17172","date":"2023-12-28","proceeding":null,"authors":["Jiasen Lu","Christopher Clark","Sangho Lee","Zichen Zhang","Savya Khosla","Ryan Marten","Derek Hoiem","Aniruddha Kembhavi"],"abstract":"We present Unified-IO 2, the first autoregressive multimodal model that is capable of understanding and generating image, text, audio, and action. To unify different modalities, we tokenize inputs and outputs -- images, text, audio, action, bounding boxes, etc., into a shared semantic space and then process them with a single encoder-decoder transformer model. Since training with such diverse modalities is challenging, we propose various architectural improvements to stabilize model training. We train our model from scratch on a large multimodal pre-training corpus from diverse sources with a multimodal mixture of denoisers objective. To learn an expansive set of skills, such as following multimodal instructions, we construct and finetune on an ensemble of 120 datasets with prompts and augmentations. With a single unified model, Unified-IO 2 achieves state-of-the-art performance on the GRIT benchmark and strong results in more than 35 benchmarks, including image generation and understanding, natural language understanding, video and audio understanding, and robotic manipulation. We release all our models to the research community.","url_abs":"https://arxiv.org/abs/2312.17172v1","url_pdf":"https://arxiv.org/pdf/2312.17172v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"unified-io-2-scaling-autoregressive","repo_url":"https://github.com/allenai/unified-io-2","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"jax","reach":null}],"tasks":[{"task_slug":"decoder","task_name":"Decoder"},{"task_slug":"image-generation","task_name":"Image Generation"},{"task_slug":"natural-language-understanding","task_name":"Natural Language Understanding"}],"methods":[{"method_slug":"set","method_name":"SET"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/2312.17172","atlas_url":"https://app.syntology.ai/?focus=2312.17172","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2312.17172"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-25T09:33:49+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/allenai/unified-io-2","reach":null}],"summary":{"ran_fixture":2},"by_repo_kind":{"official":{"samples":2,"ran":2,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"f05c7ba7b91759e2","entry":"draw_bboxes","repo":"allenai/unified-io-2","repo_kind":"official","path":"t5x/examples/unified_io/scripts/dataset_visualize.py","file_url":"https://github.com/allenai/unified-io-2/blob/HEAD/t5x/examples/unified_io/scripts/dataset_visualize.py","link_basis":"first_harvest_node","language":"python","status":"ran_fixture","verification_level":1,"contract_check":"RAISES","metamorphic_tier":"invariant","behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"f05c7ba7b91759e2"}},{"code_sha256_prefix":"8c97e690db5c9c8d","entry":"mask_image","repo":"allenai/unified-io-2","repo_kind":"official","path":"t5x/examples/unified_io/scripts/dataset_visualize.py","file_url":"https://github.com/allenai/unified-io-2/blob/HEAD/t5x/examples/unified_io/scripts/dataset_visualize.py","link_basis":"first_harvest_node","language":"python","status":"ran_fixture","verification_level":1,"contract_check":"RAISES","metamorphic_tier":"invariant","behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"8c97e690db5c9c8d"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}