{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/generalized-decoding-for-pixel-image-and","title":"Generalized Decoding for Pixel, Image, and Language","arxiv_id":"2212.11270","date":"2022-12-21","proceeding":"CVPR 2023 1","authors":["Xueyan Zou","Zi-Yi Dou","Jianwei Yang","Zhe Gan","Linjie Li","Chunyuan Li","Xiyang Dai","Harkirat Behl","JianFeng Wang","Lu Yuan","Nanyun Peng","Lijuan Wang","Yong Jae Lee","Jianfeng Gao"],"abstract":"We present X-Decoder, a generalized decoding model that can predict pixel-level segmentation and language tokens seamlessly. X-Decodert takes as input two types of queries: (i) generic non-semantic queries and (ii) semantic queries induced from text inputs, to decode different pixel-level and token-level outputs in the same semantic space. With such a novel design, X-Decoder is the first work that provides a unified way to support all types of image segmentation and a variety of vision-language (VL) tasks. Further, our design enables seamless interactions across tasks at different granularities and brings mutual benefits by learning a common and rich pixel-level visual-semantic understanding space, without any pseudo-labeling. After pretraining on a mixed set of a limited amount of segmentation data and millions of image-text pairs, X-Decoder exhibits strong transferability to a wide range of downstream tasks in both zero-shot and finetuning settings. Notably, it achieves (1) state-of-the-art results on open-vocabulary segmentation and referring segmentation on eight datasets; (2) better or competitive finetuned performance to other generalist and specialist models on segmentation and VL tasks; and (3) flexibility for efficient finetuning and novel task composition (e.g., referring captioning and image editing). Code, demo, video, and visualization are available at https://x-decoder-vl.github.io.","url_abs":"https://arxiv.org/abs/2212.11270v1","url_pdf":"https://arxiv.org/pdf/2212.11270v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"generalized-decoding-for-pixel-image-and","repo_url":"https://github.com/microsoft/X-Decoder","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"Apache-2.0"}}],"tasks":[{"task_slug":"decoder","task_name":"Decoder"},{"task_slug":"image-segmentation","task_name":"Image Segmentation"},{"task_slug":"instance-segmentation","task_name":"Instance Segmentation"},{"task_slug":"panoptic-segmentation","task_name":"Panoptic Segmentation"},{"task_slug":"referring-expression-segmentation","task_name":"Referring Expression Segmentation"},{"task_slug":"segmentation","task_name":"Segmentation"},{"task_slug":"semantic-segmentation","task_name":"Semantic Segmentation"},{"task_slug":"zero-shot-segmentation","task_name":"Zero Shot Segmentation"}],"methods":[],"datasets_introduced":[{"slug":"segmentation-in-the-wild","name":"Segmentation in the Wild","full_name":"Segmentation in the Wild"}],"methods_introduced":[],"results":[{"leaderboard":"/sota/instance-segmentation-on-ade20k-val","task":"Instance Segmentation","dataset":"ADE20K val","model":"X-Decoder (Davit-d5, Deform, single-scale, 1280x1280)","rank_in_archive_order":5,"of":14,"metrics":{"AP":"38.7","APL":"59.6","APM":"43.3","APS":"18.9"},"uses_additional_data":true},{"leaderboard":"/sota/instance-segmentation-on-ade20k-val","task":"Instance Segmentation","dataset":"ADE20K val","model":"X-Decoder (L)","rank_in_archive_order":9,"of":14,"metrics":{"AP":"35.8"},"uses_additional_data":true},{"leaderboard":"/sota/panoptic-segmentation-on-ade20k-val","task":"Panoptic Segmentation","dataset":"ADE20K val","model":"X-Decoder (Davit-d5, Deform, single-scale, 1280x1280)","rank_in_archive_order":6,"of":25,"metrics":{"AP":"38.7","PQ":"52.4","mIoU":"59.1"},"uses_additional_data":true},{"leaderboard":"/sota/panoptic-segmentation-on-ade20k-val","task":"Panoptic Segmentation","dataset":"ADE20K val","model":"X-Decoder (L)","rank_in_archive_order":15,"of":25,"metrics":{"AP":"35.8","PQ":"49.6","mIoU":"58.1"},"uses_additional_data":true},{"leaderboard":"/sota/referring-expression-segmentation-on-refcocog","task":"Referring Expression Segmentation","dataset":"RefCOCOg-val","model":"X-Decoder (Davit-d5)","rank_in_archive_order":17,"of":23,"metrics":{"Overall IoU":"64.6"},"uses_additional_data":true},{"leaderboard":"/sota/zero-shot-segmentation-on-segmentation-in-the","task":"Zero Shot Segmentation","dataset":"Segmentation in the Wild","model":"SGinW_Team (X-Decoder-L)","rank_in_archive_order":9,"of":12,"metrics":{"Mean AP":"32.2"},"uses_additional_data":false},{"leaderboard":"/sota/zero-shot-segmentation-on-segmentation-in-the","task":"Zero Shot Segmentation","dataset":"Segmentation in the Wild","model":"SGinW_Team (X-Decoder-B)","rank_in_archive_order":10,"of":12,"metrics":{"Mean AP":"27.7"},"uses_additional_data":false},{"leaderboard":"/sota/zero-shot-segmentation-on-segmentation-in-the","task":"Zero Shot Segmentation","dataset":"Segmentation in the Wild","model":"SGinW_Team (X-Decoder-L-IN21K)","rank_in_archive_order":11,"of":12,"metrics":{"Mean AP":"26.6"},"uses_additional_data":false},{"leaderboard":"/sota/zero-shot-segmentation-on-segmentation-in-the","task":"Zero Shot Segmentation","dataset":"Segmentation in the Wild","model":"SGinW_Team (X-Decoder-T)","rank_in_archive_order":12,"of":12,"metrics":{"Mean AP":"22.6"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2212.11270","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}