{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/improving-weakly-supervised-visual-grounding","title":"Improving Weakly Supervised Visual Grounding by Contrastive Knowledge Distillation","arxiv_id":"2007.01951","date":"2020-07-03","proceeding":"CVPR 2021 1","authors":["Liwei Wang","Jing Huang","Yin Li","Kun Xu","Zhengyuan Yang","Dong Yu"],"abstract":"Weakly supervised phrase grounding aims at learning region-phrase correspondences using only image-sentence pairs. A major challenge thus lies in the missing links between image regions and sentence phrases during training. To address this challenge, we leverage a generic object detector at training time, and propose a contrastive learning framework that accounts for both region-phrase and image-sentence matching. Our core innovation is the learning of a region-phrase score function, based on which an image-sentence score function is further constructed. Importantly, our region-phrase score function is learned by distilling from soft matching scores between the detected object names and candidate phrases within an image-sentence pair, while the image-sentence score function is supervised by ground-truth image-sentence pairs. The design of such score functions removes the need of object detection at test time, thereby significantly reducing the inference cost. Without bells and whistles, our approach achieves state-of-the-art results on visual phrase grounding, surpassing previous methods that require expensive object detectors at test time.","url_abs":"https://arxiv.org/abs/2007.01951v2","url_pdf":"https://arxiv.org/pdf/2007.01951v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"improving-weakly-supervised-visual-grounding","repo_url":"https://github.com/jhuang81/weak-sup-visual-grounding","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"contrastive-learning","task_name":"Contrastive Learning"},{"task_slug":"knowledge-distillation","task_name":"Knowledge Distillation"},{"task_slug":"object","task_name":"Object"},{"task_slug":"object-detection","task_name":"Object Detection"},{"task_slug":"phrase-grounding","task_name":"Phrase Grounding"},{"task_slug":"sentence","task_name":"Sentence"},{"task_slug":"visual-grounding","task_name":"Visual Grounding"},{"task_slug":"object-detection-1","task_name":"object-detection"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2007.01951","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2007.01951"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/jhuang81/weak-sup-visual-grounding","reach":{"status":"ok","spdx":"MIT"}}],"summary":{"unverified":3},"by_repo_kind":{"official":{"samples":3,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"c7e6511c79500c56","entry":"cpc_loss","repo":"jhuang81/weak-sup-visual-grounding","repo_kind":"official","path":"nce_distill_model.py","file_url":"https://github.com/jhuang81/weak-sup-visual-grounding/blob/HEAD/nce_distill_model.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"c7e6511c79500c56"}},{"code_sha256_prefix":"dd0e8bb108fe0b39","entry":"feedforward_net","repo":"jhuang81/weak-sup-visual-grounding","repo_kind":"official","path":"nce_distill_model.py","file_url":"https://github.com/jhuang81/weak-sup-visual-grounding/blob/HEAD/nce_distill_model.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"dd0e8bb108fe0b39"}},{"code_sha256_prefix":"0249f7e71be8ec2b","entry":"parse_pbtxt","repo":"jhuang81/weak-sup-visual-grounding","repo_kind":"official","path":"oiv2_classes.py","file_url":"https://github.com/jhuang81/weak-sup-visual-grounding/blob/HEAD/oiv2_classes.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"0249f7e71be8ec2b"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}