{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/generative-prompt-model-for-weakly-supervised","title":"Generative Prompt Model for Weakly Supervised Object Localization","arxiv_id":"2307.09756","date":"2023-07-19","proceeding":"ICCV 2023 1","authors":["Yuzhong Zhao","Qixiang Ye","Weijia Wu","Chunhua Shen","Fang Wan"],"abstract":"Weakly supervised object localization (WSOL) remains challenging when learning object localization models from image category labels. Conventional methods that discriminatively train activation models ignore representative yet less discriminative object parts. In this study, we propose a generative prompt model (GenPromp), defining the first generative pipeline to localize less discriminative object parts by formulating WSOL as a conditional image denoising procedure. During training, GenPromp converts image category labels to learnable prompt embeddings which are fed to a generative model to conditionally recover the input image with noise and learn representative embeddings. During inference, enPromp combines the representative embeddings with discriminative embeddings (queried from an off-the-shelf vision-language model) for both representative and discriminative capacity. The combined embeddings are finally used to generate multi-scale high-quality attention maps, which facilitate localizing full object extent. Experiments on CUB-200-2011 and ILSVRC show that GenPromp respectively outperforms the best discriminative models by 5.2% and 5.6% (Top-1 Loc), setting a solid baseline for WSOL with the generative model. Code is available at https://github.com/callsys/GenPromp.","url_abs":"https://arxiv.org/abs/2307.09756v1","url_pdf":"https://arxiv.org/pdf/2307.09756v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"generative-prompt-model-for-weakly-supervised","repo_url":"https://github.com/callsys/genpromp","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"denoising","task_name":"Denoising"},{"task_slug":"image-denoising","task_name":"Image Denoising"},{"task_slug":"language-modeling","task_name":"Language Modeling"},{"task_slug":"language-modelling","task_name":"Language Modelling"},{"task_slug":"object","task_name":"Object"},{"task_slug":"object-localization","task_name":"Object Localization"},{"task_slug":"weakly-supervised-object-localization","task_name":"Weakly-Supervised Object Localization"},{"task_slug":"model","task_name":"model"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/weakly-supervised-object-localization-on-cub-2","task":"Weakly-Supervised Object Localization","dataset":"CUB-200-2011","model":"Stable diffusion","rank_in_archive_order":3,"of":3,"metrics":{"GT-known localization accuracy":"98.0","Top-1 Localization Accuracy":"87.0"},"uses_additional_data":true},{"leaderboard":"/sota/weakly-supervised-object-localization-on-cub","task":"Weakly-Supervised Object Localization","dataset":"CUB-200-2011","model":"GenPromp","rank_in_archive_order":5,"of":10,"metrics":{"Top-1 Localization Accuracy":"87.0"},"uses_additional_data":true},{"leaderboard":"/sota/weakly-supervised-object-localization-on-2","task":"Weakly-Supervised Object Localization","dataset":"ImageNet","model":"Stable diffusion","rank_in_archive_order":1,"of":6,"metrics":{"GT-known localization accuracy":"75.0","Top-1 Localization Accuracy":"65.2"},"uses_additional_data":true}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2307.09756","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2307.09756"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"deterministic:regex_extraction","url":"https://github.com/callsys/GenPromp","reach":null}],"summary":{"ran":1},"by_repo_kind":{"official":{"samples":1,"ran":1,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"1fc03898fe70708a","entry":"AttentionStore","repo":"callsys/GenPromp","repo_kind":"official","path":"models/attn.py","file_url":"https://github.com/callsys/GenPromp/blob/HEAD/models/attn.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"1fc03898fe70708a"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}