{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/unleashing-text-to-image-diffusion-models-for-1","title":"Unleashing Text-to-Image Diffusion Models for Visual Perception","arxiv_id":"2303.02153","date":"2023-03-03","proceeding":"ICCV 2023 1","authors":["Wenliang Zhao","Yongming Rao","Zuyan Liu","Benlin Liu","Jie zhou","Jiwen Lu"],"abstract":"Diffusion models (DMs) have become the new trend of generative models and have demonstrated a powerful ability of conditional synthesis. Among those, text-to-image diffusion models pre-trained on large-scale image-text pairs are highly controllable by customizable prompts. Unlike the unconditional generative models that focus on low-level attributes and details, text-to-image diffusion models contain more high-level knowledge thanks to the vision-language pre-training. In this paper, we propose VPD (Visual Perception with a pre-trained Diffusion model), a new framework that exploits the semantic information of a pre-trained text-to-image diffusion model in visual perception tasks. Instead of using the pre-trained denoising autoencoder in a diffusion-based pipeline, we simply use it as a backbone and aim to study how to take full advantage of the learned knowledge. Specifically, we prompt the denoising decoder with proper textual inputs and refine the text features with an adapter, leading to a better alignment to the pre-trained stage and making the visual contents interact with the text prompts. We also propose to utilize the cross-attention maps between the visual features and the text features to provide explicit guidance. Compared with other pre-training methods, we show that vision-language pre-trained diffusion models can be faster adapted to downstream visual perception tasks using the proposed VPD. Extensive experiments on semantic segmentation, referring image segmentation and depth estimation demonstrates the effectiveness of our method. Notably, VPD attains 0.254 RMSE on NYUv2 depth estimation and 73.3% oIoU on RefCOCO-val referring image segmentation, establishing new records on these two benchmarks. Code is available at https://github.com/wl-zhao/VPD","url_abs":"https://arxiv.org/abs/2303.02153v1","url_pdf":"https://arxiv.org/pdf/2303.02153v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"unleashing-text-to-image-diffusion-models-for-1","repo_url":"https://github.com/wl-zhao/VPD","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"unleashing-text-to-image-diffusion-models-for-1","repo_url":"https://github.com/open-mmlab/mmsegmentation/tree/main/configs/vpd","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"denoising","task_name":"Denoising"},{"task_slug":"depth-estimation","task_name":"Depth Estimation"},{"task_slug":"image-segmentation","task_name":"Image Segmentation"},{"task_slug":"monocular-depth-estimation","task_name":"Monocular Depth Estimation"},{"task_slug":"referring-expression-segmentation","task_name":"Referring Expression Segmentation"},{"task_slug":"segmentation","task_name":"Segmentation"},{"task_slug":"semantic-segmentation","task_name":"Semantic Segmentation"}],"methods":[{"method_slug":"denoising-autoencoder","method_name":"Denoising Autoencoder"},{"method_slug":"diffusion","method_name":"Diffusion"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/monocular-depth-estimation-on-nyu-depth-v2","task":"Monocular Depth Estimation","dataset":"NYU-Depth V2","model":"VPD","rank_in_archive_order":21,"of":85,"metrics":{"Delta < 1.25":"0.964","Delta < 1.25^2":"0.995","Delta < 1.25^3":"0.999","RMSE":"0.254","absolute relative error":"0.069","log 10":"0.030"},"uses_additional_data":false},{"leaderboard":"/sota/referring-expression-segmentation-on-refcoco","task":"Referring Expression Segmentation","dataset":"RefCoCo val","model":"VPD","rank_in_archive_order":21,"of":37,"metrics":{"Overall IoU":"73.25"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2303.02153","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2303.02153"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/open-mmlab/mmsegmentation/tree/main/configs/vpd","reach":null},{"provenance":"deterministic:regex_extraction","url":"https://github.com/wl-zhao/VPD","reach":{"status":"ok","spdx":"MIT"}}],"summary":{"ran_violates":2,"ran":1,"ran_draft_wrong":1,"unverified":1},"by_repo_kind":{"official":{"samples":5,"ran":4,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"424012cb37b31172","entry":"default","repo":"wl-zhao/VPD","repo_kind":"official","path":"vpd/models.py","file_url":"https://github.com/wl-zhao/VPD/blob/HEAD/vpd/models.py","link_basis":"harvester_set","language":"python","status":"ran_violates","verification_level":1,"contract_check":"VIOLATES","metamorphic_tier":"well_formed","behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"424012cb37b31172"}},{"code_sha256_prefix":"aa5486a3650902d8","entry":"exists","repo":"wl-zhao/VPD","repo_kind":"official","path":"vpd/models.py","file_url":"https://github.com/wl-zhao/VPD/blob/HEAD/vpd/models.py","link_basis":"harvester_set","language":"python","status":"ran_violates","verification_level":1,"contract_check":"VIOLATES","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"aa5486a3650902d8"}},{"code_sha256_prefix":"6907f72868217c02","entry":"get_num_layer_for_swin","repo":"wl-zhao/VPD","repo_kind":"official","path":"depth/models_depth/optimizer.py","file_url":"https://github.com/wl-zhao/VPD/blob/HEAD/depth/models_depth/optimizer.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"6907f72868217c02"}},{"code_sha256_prefix":"9a299fe5ae09e407","entry":"uniq","repo":"wl-zhao/VPD","repo_kind":"official","path":"vpd/models.py","file_url":"https://github.com/wl-zhao/VPD/blob/HEAD/vpd/models.py","link_basis":"harvester_set","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"9a299fe5ae09e407"}},{"code_sha256_prefix":"b700fb0a3496b255","entry":"cosine_scheduler","repo":"wl-zhao/VPD","repo_kind":"official","path":"depth/utils.py","file_url":"https://github.com/wl-zhao/VPD/blob/HEAD/depth/utils.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"b700fb0a3496b255"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}