{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/unified-io-a-unified-model-for-vision","title":"Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks","arxiv_id":"2206.08916","date":"2022-06-17","proceeding":null,"authors":["Jiasen Lu","Christopher Clark","Rowan Zellers","Roozbeh Mottaghi","Aniruddha Kembhavi"],"abstract":"We propose Unified-IO, a model that performs a large variety of AI tasks spanning classical computer vision tasks, including pose estimation, object detection, depth estimation and image generation, vision-and-language tasks such as region captioning and referring expression, to natural language processing tasks such as question answering and paraphrasing. Developing a single unified model for such a large variety of tasks poses unique challenges due to the heterogeneous inputs and outputs pertaining to each task, including RGB images, per-pixel maps, binary masks, bounding boxes, and language. We achieve this unification by homogenizing every supported input and output into a sequence of discrete vocabulary tokens. This common representation across all tasks allows us to train a single transformer-based architecture, jointly on over 90 diverse datasets in the vision and language fields. Unified-IO is the first model capable of performing all 7 tasks on the GRIT benchmark and produces strong results across 16 diverse benchmarks like NYUv2-Depth, ImageNet, VQA2.0, OK-VQA, Swig, VizWizGround, BoolQ, and SciTail, with no task-specific fine-tuning. Code and demos for Unified-IO are available at: https://unified-io.allenai.org.","url_abs":"https://arxiv.org/abs/2206.08916v2","url_pdf":"https://arxiv.org/pdf/2206.08916v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"depth-estimation","task_name":"Depth Estimation"},{"task_slug":"image-generation","task_name":"Image Generation"},{"task_slug":"keypoint-estimation","task_name":"Keypoint Estimation"},{"task_slug":"object-categorization","task_name":"Object Categorization"},{"task_slug":"object-detection","task_name":"Object Detection"},{"task_slug":"object-localization","task_name":"Object Localization"},{"task_slug":"object-segmentation","task_name":"Object Segmentation"},{"task_slug":"pose-estimation","task_name":"Pose Estimation"},{"task_slug":"question-answering","task_name":"Question Answering"},{"task_slug":"referring-expression","task_name":"Referring Expression"},{"task_slug":"referring-expression-comprehension","task_name":"Referring Expression Comprehension"},{"task_slug":"surface-normal-estimation","task_name":"Surface Normal Estimation"},{"task_slug":"visual-question-answering","task_name":"Visual Question Answering (VQA)"},{"task_slug":"object-detection-1","task_name":"object-detection"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/object-categorization-on-grit","task":"Object Categorization","dataset":"GRIT","model":"Unified-IOXL","rank_in_archive_order":1,"of":4,"metrics":{"Categorization (ablation)":"61.7","Categorization (test)":"60.8"},"uses_additional_data":false},{"leaderboard":"/sota/object-localization-on-grit","task":"Object Localization","dataset":"GRIT","model":"Unified-IOXL","rank_in_archive_order":1,"of":3,"metrics":{"Localization (ablation)":"67.0","Localization (test)":"67.1"},"uses_additional_data":false},{"leaderboard":"/sota/object-segmentation-on-grit","task":"Object Segmentation","dataset":"GRIT","model":"Unified-IOXL","rank_in_archive_order":1,"of":2,"metrics":{"Segmentation (ablation)":"56.3","Segmentation (test)":"56.5"},"uses_additional_data":false},{"leaderboard":"/sota/visual-question-answering-on-grit","task":"Visual Question Answering (VQA)","dataset":"GRIT","model":"Unified-IOXL","rank_in_archive_order":1,"of":2,"metrics":{"VQA (ablation)":"74.5","VQA (test)":"74.5"},"uses_additional_data":false}],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=2206.08916","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}