Papers › Visual Translation Embedding Network for Visual Relation Detection

Visual Translation Embedding Network for Visual Relation Detection

27 Feb 2017CVPR 2017 7arXiv:1702.08319archive 2025-07-28

Hanwang Zhang, Zawlin Kyaw, Shih-Fu Chang, Tat-Seng Chua

Visual relations, such as "person ride bike" and "bike next to car", offer a comprehensive scene understanding of an image, and have already shown their great utility in connecting computer vision and natural language. However, due to the challenging combinatorial complexity of modeling subject-predicate-object relation triplets, very little work has been done to localize and predict visual relations. Inspired by the recent advances in relational representation learning of knowledge bases and convolutional object detection networks, we propose a Visual Translation Embedding network (VTransE) for visual relation detection. VTransE places objects in a low-dimensional relation space where a relation can be modeled as a simple vector translation, i.e., subject + predicate ≈ object. We propose a novel feature extraction layer that enables object-relation knowledge transfer in a fully-convolutional fashion that supports training and inference in a single forward/backward pass. To the best of our knowledge, VTransE is the first end-to-end relation detection network. We demonstrate the effectiveness of VTransE over other state-of-the-art methods on two large-scale datasets: Visual Relationship and Visual Genome. Note that even though VTransE is a purely visual model, it is still competitive to the Lu's multi-modal model with language priors.

PaperPDFConference PDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

zawlin/cvpr17_vtranse caffe2NOASSERTION report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

ObjectObject DetectionRepresentation LearningScene UnderstandingTransfer LearningTranslationobject-detection

1 archive task tag without a task page not shown.

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Visual Relationship Detection VRD Phrase Detection Zhang et. al [[Hanwang Zhang2017]] R@100 22.42 #5 of 7 Archive leaderboard report
Visual Relationship Detection VRD Phrase Detection Zhang et. al [[Hanwang Zhang2017]] R@50 19.42 #5 of 7 Archive leaderboard report
Visual Relationship Detection VRD Predicate Detection Zhang et. al [[Hanwang Zhang2017]] R@100 44.76 #7 of 7 Archive leaderboard report
Visual Relationship Detection VRD Predicate Detection Zhang et. al [[Hanwang Zhang2017]] R@50 44.76 #7 of 7 Archive leaderboard report
Visual Relationship Detection VRD Relationship Detection Zhang et. al [[Hanwang Zhang2017]] R@100 15.20 #7 of 8 Archive leaderboard report
Visual Relationship Detection VRD Relationship Detection Zhang et. al [[Hanwang Zhang2017]] R@50 14.07 #7 of 8 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections