Papers › What and When to Look?: Temporal Span Proposal Network for Video Relation Detection
What and When to Look?: Temporal Span Proposal Network for Video Relation Detection
Sangmin Woo, Junhyug Noh, Kangil Kim
Identifying relations between objects is central to understanding the scene. While several works have been proposed for relation modeling in the image domain, there have been many constraints in the video domain due to challenging dynamics of spatio-temporal interactions (e.g., between which objects are there an interaction? when do relations start and end?). To date, two representative methods have been proposed to tackle Video Visual Relation Detection (VidVRD): segment-based and window-based. We first point out limitations of these methods and propose a novel approach named Temporal Span Proposal Network (TSPN). TSPN tells what to look: it sparsifies relation search space by scoring relationness of object pair, i.e., measuring how probable a relation exist. TSPN tells when to look: it simultaneously predicts start-end timestamps (i.e., temporal spans) and categories of the all possible relations by utilizing full video context. These two designs enable a win-win scenario: it accelerates training by 2X or more than existing methods and achieves competitive performance on two VidVRD benchmarks (ImageNet-VidVDR and VidOR). Moreover, comprehensive ablative experiments demonstrate the effectiveness of our approach. Codes are available at https://github.com/sangminwoo/Temporal-Span-Proposal-Network-VidVRD.
Code
Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
1 archive task tag without a task page not shown.
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| Video Visual Relation Detection | ImageNet-VidVRD | TSPN | Recall@100 | 14.13 | #2 of 2 | Archive leaderboard | report |
| Video Visual Relation Detection | ImageNet-VidVRD | TSPN | Recall@50 | 11.56 | #2 of 2 | Archive leaderboard | report |
| Video Visual Relation Detection | ImageNet-VidVRD | TSPN | mAP | 18.9 | #2 of 2 | Archive leaderboard | report |
| Video Visual Relation Detection | VidOR | TSPN | Recall@100 | 10.71 | #2 of 2 | Archive leaderboard | report |
| Video Visual Relation Detection | VidOR | TSPN | Recall@50 | 9.33 | #2 of 2 | Archive leaderboard | report |
| Video Visual Relation Detection | VidOR | TSPN | mAP | 7.61 | #2 of 2 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections