Papers › Dynamic Scene Understanding from Vision-Language Representations

Dynamic Scene Understanding from Vision-Language Representations

20 Jan 2025arXiv:2501.11653archive 2025-07-28

Shahaf Pruss, Morris Alper, Hadar Averbuch-Elor

Images depicting complex, dynamic scenes are challenging to parse automatically, requiring both high-level comprehension of the overall situation and fine-grained identification of participating entities and their interactions. Current approaches use distinct methods tailored to sub-tasks such as Situation Recognition and detection of Human-Human and Human-Object Interactions. However, recent advances in image understanding have often leveraged web-scale vision-language (V&L) representations to obviate task-specific engineering. In this work, we propose a framework for dynamic scene understanding tasks by leveraging knowledge from modern, frozen V&L representations. By framing these tasks in a generic manner - as predicting and parsing structured text, or by directly concatenating representations to the input of existing models - we achieve state-of-the-art results while using a minimal number of trainable parameters relative to existing approaches. Moreover, our analysis of dynamic knowledge of these representations shows that recent, more powerful representations effectively encode dynamic scene semantics, making this approach newly possible.

PaperPDF

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Grounded Situation RecognitionHuman Interaction RecognitionHuman-Human Interaction RecognitionHuman-Object Interaction DetectionScene UnderstandingSituation Recognition

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Grounded Situation Recognition SWiG Ours (CoFormer+) Top-1 Verb 58.88 #1 of 13 Archive leaderboard report
Grounded Situation Recognition SWiG Ours (CoFormer+) Top-1 Verb & Grounded-Value 41.28 #1 of 13 Archive leaderboard report
Grounded Situation Recognition SWiG Ours (CoFormer+) Top-1 Verb & Value 51.10 #1 of 13 Archive leaderboard report
Grounded Situation Recognition SWiG Ours (CoFormer+) Top-5 Verbs & Grounded-Value 58.23 #1 of 13 Archive leaderboard report
Human-Object Interaction Detection HICO-DET Ours (PViC+) mAP 46.49 #1 of 55 Archive leaderboard report
Situation Recognition imSitu Ours Top-1 Verb 58.88 #1 of 13 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections