Papers › Visual Translation Embedding Network for Visual Relation Detection
Visual Translation Embedding Network for Visual Relation Detection
Hanwang Zhang, Zawlin Kyaw, Shih-Fu Chang, Tat-Seng Chua
Visual relations, such as "person ride bike" and "bike next to car", offer a comprehensive scene understanding of an image, and have already shown their great utility in connecting computer vision and natural language. However, due to the challenging combinatorial complexity of modeling subject-predicate-object relation triplets, very little work has been done to localize and predict visual relations. Inspired by the recent advances in relational representation learning of knowledge bases and convolutional object detection networks, we propose a Visual Translation Embedding network (VTransE) for visual relation detection. VTransE places objects in a low-dimensional relation space where a relation can be modeled as a simple vector translation, i.e., subject + predicate ≈ object. We propose a novel feature extraction layer that enables object-relation knowledge transfer in a fully-convolutional fashion that supports training and inference in a single forward/backward pass. To the best of our knowledge, VTransE is the first end-to-end relation detection network. We demonstrate the effectiveness of VTransE over other state-of-the-art methods on two large-scale datasets: Visual Relationship and Visual Genome. Note that even though VTransE is a purely visual model, it is still competitive to the Lu's multi-modal model with language priors.
In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.
Code
Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
1 archive task tag without a task page not shown.
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| Visual Relationship Detection | VRD Phrase Detection | Zhang et. al [[Hanwang Zhang2017]] | R@100 | 22.42 | #5 of 7 | Archive leaderboard | report |
| Visual Relationship Detection | VRD Phrase Detection | Zhang et. al [[Hanwang Zhang2017]] | R@50 | 19.42 | #5 of 7 | Archive leaderboard | report |
| Visual Relationship Detection | VRD Predicate Detection | Zhang et. al [[Hanwang Zhang2017]] | R@100 | 44.76 | #7 of 7 | Archive leaderboard | report |
| Visual Relationship Detection | VRD Predicate Detection | Zhang et. al [[Hanwang Zhang2017]] | R@50 | 44.76 | #7 of 7 | Archive leaderboard | report |
| Visual Relationship Detection | VRD Relationship Detection | Zhang et. al [[Hanwang Zhang2017]] | R@100 | 15.20 | #7 of 8 | Archive leaderboard | report |
| Visual Relationship Detection | VRD Relationship Detection | Zhang et. al [[Hanwang Zhang2017]] | R@50 | 14.07 | #7 of 8 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections