Browse State-of-the-Art › Video Visual Relation Detection
Video Visual Relation Detection
9 papers with code · 2 benchmarks · 2 datasets archive 2025-07-28
Video Visual Relation Detection (VidVRD) aims to detect instances of visual relations of interest in a video, where a visual relation instance is represented by a relation triplet with the trajectories of the subject and object. As compared to still images, videos provide a more natural set of features for detecting visual relations, such as the dynamic relations like “A-follow-B” and “A-towards-B”, and temporally changing relations like “A-chase-B” followed by “A-hold-B”. Yet, VidVRD is technically more challenging than ImgVRD due to the difficulties in accurate object tracking and diverse relation appearances in the video domain.
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
2 leaderboard tables shown for this task, 2 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| ImageNet-VidVRD (2 rows) | Social Fabric | Social Fabric: Tubelet Compositions for Video Relation Detection | code | — | Compare |
| VidOR (2 rows) | Social Fabric | Social Fabric: Tubelet Compositions for Video Relation Detection | code | — | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
2 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
9 shown of 9 papers with code (15 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
26 Jul 2021 2 repositories listed Syntology ran 7 of 14 samples · 7 unverified · 5 pointer-only (licence)Compared to the task of scene graph generation from images, it is more challenging because of the dynamic relationships between objects and the temporal dependencies between frames allowing for a richer semantic…
-
18 Aug 2024 1 repository listedVideo Visual Relation Detection (VidVRD) focuses on understanding how entities interact over time and space in videos, a key step for gaining deeper insights into video scenes beyond basic visual tasks.
-
6 Apr 2024 1 repository listedWe hope that SportsHHI can stimulate research on human interaction understanding in videos and promote the development of spatio-temporal context modeling techniques in video visual relation detection.
-
1 Feb 2023 1 repository listed Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)Without bells and whistles, our RePro achieves a new state-of-the-art performance on two VidVRD benchmarks of not only the base training object and predicate categories, but also the unseen ones.
-
19 Aug 2021 1 repository listedVideo Visual Relation Detection (VidVRD), has received significant attention of our community over recent years.
-
18 Aug 2021 1 repository listedWe also propose Social Fabric: an encoding that represents a pair of object tubelets as a composition of interaction primitives.
-
15 Jul 2021 1 repository listedTSPN tells when to look: it simultaneously predicts start-end timestamps (i.
-
17 Dec 2020 1 repository listedAnalyzing the interactions between humans and objects from a video includes identification of the relationships between humans and the objects present in the video.
-
25 Mar 2019 1 repository listedVisual relationship reasoning is a crucial yet challenging task for understanding rich interactions across visual concepts.
Syntology lines on 2 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections