{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/weakly-supervised-learning-of-visual","title":"Weakly-supervised learning of visual relations","arxiv_id":"1707.09472","date":"2017-07-29","proceeding":"ICCV 2017 10","authors":["Julia Peyre","Ivan Laptev","Cordelia Schmid","Josef Sivic"],"abstract":"This paper introduces a novel approach for modeling visual relations between\npairs of objects. We call relation a triplet of the form (subject, predicate,\nobject) where the predicate is typically a preposition (eg. 'under', 'in front\nof') or a verb ('hold', 'ride') that links a pair of objects (subject, object).\nLearning such relations is challenging as the objects have different spatial\nconfigurations and appearances depending on the relation in which they occur.\nAnother major challenge comes from the difficulty to get annotations,\nespecially at box-level, for all possible triplets, which makes both learning\nand evaluation difficult. The contributions of this paper are threefold. First,\nwe design strong yet flexible visual features that encode the appearance and\nspatial configuration for pairs of objects. Second, we propose a\nweakly-supervised discriminative clustering model to learn relations from\nimage-level labels only. Third we introduce a new challenging dataset of\nunusual relations (UnRel) together with an exhaustive annotation, that enables\naccurate evaluation of visual relation retrieval. We show experimentally that\nour model results in state-of-the-art results on the visual relationship\ndataset significantly improving performance on previously unseen relations\n(zero-shot learning), and confirm this observation on our newly introduced\nUnRel dataset.","url_abs":"http://arxiv.org/abs/1707.09472v1","url_pdf":"http://arxiv.org/pdf/1707.09472v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"clustering","task_name":"Clustering"},{"task_slug":null,"task_name":"Relation"},{"task_slug":"retrieval","task_name":"Retrieval"},{"task_slug":null,"task_name":"Triplet"},{"task_slug":"weakly-supervised-learning","task_name":"Weakly-supervised Learning"},{"task_slug":"zero-shot-learning","task_name":"Zero-Shot Learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/visual-relationship-detection-on-vrd-phrase","task":"Visual Relationship Detection","dataset":"VRD Phrase Detection","model":"Peyre et. al [[Peyre et al.2017]]","rank_in_archive_order":6,"of":7,"metrics":{"R@100":"19.5","R@50":"17.9"},"uses_additional_data":false},{"leaderboard":"/sota/visual-relationship-detection-on-vrd","task":"Visual Relationship Detection","dataset":"VRD Predicate Detection","model":"Peyre et. al [[Peyre et al.2017]]","rank_in_archive_order":5,"of":7,"metrics":{"R@100":"52.6","R@50":"52.6"},"uses_additional_data":false},{"leaderboard":"/sota/visual-relationship-detection-on-vrd-1","task":"Visual Relationship Detection","dataset":"VRD Relationship Detection","model":"Peyre et. al [[Peyre et al.2017]]","rank_in_archive_order":6,"of":8,"metrics":{"R@100":"17.1","R@50":"15.8"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1707.09472","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}