Papers › Visual Relationship Detection with Internal and External Linguistic Knowledge Distillation
Visual Relationship Detection with Internal and External Linguistic Knowledge Distillation
Ruichi Yu, Ang Li, Vlad I. Morariu, Larry S. Davis
Understanding visual relationships involves identifying the subject, the object, and a predicate relating them. We leverage the strong correlations between the predicate and the (subj,obj) pair (both semantically and spatially) to predict the predicates conditioned on the subjects and the objects. Modeling the three entities jointly more accurately reflects their relationships, but complicates learning since the semantic space of visual relationships is huge and the training data is limited, especially for the long-tail relationships that have few instances. To overcome this, we use knowledge of linguistic statistics to regularize visual model learning. We obtain linguistic knowledge by mining from both training annotations (internal knowledge) and publicly available text, e.g., Wikipedia (external knowledge), computing the conditional probability distribution of a predicate given a (subj,obj) pair. Then, we distill the knowledge into a deep model to achieve better generalization. Our experimental results on the Visual Relationship Detection (VRD) and Visual Genome datasets suggest that with this linguistic knowledge distillation, our model outperforms the state-of-the-art methods significantly, especially when predicting unseen relationships (e.g., recall improved from 8.45% to 19.17% on VRD zero-shot testing set).
In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.
Code
No code repository is listed for this paper in the archive or in Syntology's graph.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| Visual Relationship Detection | VRD Phrase Detection | Yu et. al [[Yu et al.2017a]] | R@100 | 29.43 | #1 of 7 | Archive leaderboard | report |
| Visual Relationship Detection | VRD Phrase Detection | Yu et. al [[Yu et al.2017a]] | R@50 | 26.32 | #1 of 7 | Archive leaderboard | report |
| Visual Relationship Detection | VRD Predicate Detection | Yu et. al [[Yu et al.2017a]] | R@100 | 94.65 | #1 of 7 | Archive leaderboard | report |
| Visual Relationship Detection | VRD Predicate Detection | Yu et. al [[Yu et al.2017a]] | R@50 | 85.64 | #1 of 7 | Archive leaderboard | report |
| Visual Relationship Detection | VRD Relationship Detection | Yu et. al [[Yu et al.2017a]] | R@100 | 31.89 | #1 of 8 | Archive leaderboard | report |
| Visual Relationship Detection | VRD Relationship Detection | Yu et. al [[Yu et al.2017a]] | R@50 | 22.68 | #1 of 8 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections