{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/visual-translation-embedding-network-for","title":"Visual Translation Embedding Network for Visual Relation Detection","arxiv_id":"1702.08319","date":"2017-02-27","proceeding":"CVPR 2017 7","authors":["Hanwang Zhang","Zawlin Kyaw","Shih-Fu Chang","Tat-Seng Chua"],"abstract":"Visual relations, such as \"person ride bike\" and \"bike next to car\", offer a\ncomprehensive scene understanding of an image, and have already shown their\ngreat utility in connecting computer vision and natural language. However, due\nto the challenging combinatorial complexity of modeling\nsubject-predicate-object relation triplets, very little work has been done to\nlocalize and predict visual relations. Inspired by the recent advances in\nrelational representation learning of knowledge bases and convolutional object\ndetection networks, we propose a Visual Translation Embedding network (VTransE)\nfor visual relation detection. VTransE places objects in a low-dimensional\nrelation space where a relation can be modeled as a simple vector translation,\ni.e., subject + predicate $\\approx$ object. We propose a novel feature\nextraction layer that enables object-relation knowledge transfer in a\nfully-convolutional fashion that supports training and inference in a single\nforward/backward pass. To the best of our knowledge, VTransE is the first\nend-to-end relation detection network. We demonstrate the effectiveness of\nVTransE over other state-of-the-art methods on two large-scale datasets: Visual\nRelationship and Visual Genome. Note that even though VTransE is a purely\nvisual model, it is still competitive to the Lu's multi-modal model with\nlanguage priors.","url_abs":"http://arxiv.org/abs/1702.08319v1","url_pdf":"http://arxiv.org/pdf/1702.08319v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"visual-translation-embedding-network-for","repo_url":"https://github.com/yangxuntu/vrd","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"tf","reach":{"status":"ok"}},{"paper_slug":"visual-translation-embedding-network-for","repo_url":"https://github.com/zawlin/cvpr17_vtranse","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"caffe2","reach":{"status":"ok","spdx":"NOASSERTION"}}],"tasks":[{"task_slug":"object","task_name":"Object"},{"task_slug":"object-detection","task_name":"Object Detection"},{"task_slug":null,"task_name":"Relation"},{"task_slug":"representation-learning","task_name":"Representation Learning"},{"task_slug":"scene-understanding","task_name":"Scene Understanding"},{"task_slug":"transfer-learning","task_name":"Transfer Learning"},{"task_slug":"translation","task_name":"Translation"},{"task_slug":"object-detection-1","task_name":"object-detection"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/visual-relationship-detection-on-vrd-phrase","task":"Visual Relationship Detection","dataset":"VRD Phrase Detection","model":"Zhang et. al [[Hanwang Zhang2017]]","rank_in_archive_order":5,"of":7,"metrics":{"R@100":"22.42","R@50":"19.42"},"uses_additional_data":false},{"leaderboard":"/sota/visual-relationship-detection-on-vrd","task":"Visual Relationship Detection","dataset":"VRD Predicate Detection","model":"Zhang et. al [[Hanwang Zhang2017]]","rank_in_archive_order":7,"of":7,"metrics":{"R@100":"44.76","R@50":"44.76"},"uses_additional_data":false},{"leaderboard":"/sota/visual-relationship-detection-on-vrd-1","task":"Visual Relationship Detection","dataset":"VRD Relationship Detection","model":"Zhang et. al [[Hanwang Zhang2017]]","rank_in_archive_order":7,"of":8,"metrics":{"R@100":"15.20","R@50":"14.07"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1702.08319","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}