{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/video-relationship-reasoning-using-gated","title":"Video Relationship Reasoning using Gated Spatio-Temporal Energy Graph","arxiv_id":"1903.10547","date":"2019-03-25","proceeding":"CVPR 2019 6","authors":["Yao-Hung Hubert Tsai","Santosh Divvala","Louis-Philippe Morency","Ruslan Salakhutdinov","Ali Farhadi"],"abstract":"Visual relationship reasoning is a crucial yet challenging task for\nunderstanding rich interactions across visual concepts. For example, a\nrelationship 'man, open, door' involves a complex relation 'open' between\nconcrete entities 'man, door'. While much of the existing work has studied this\nproblem in the context of still images, understanding visual relationships in\nvideos has received limited attention. Due to their temporal nature, videos\nenable us to model and reason about a more comprehensive set of visual\nrelationships, such as those requiring multiple (temporal) observations (e.g.,\n'man, lift up, box' vs. 'man, put down, box'), as well as relationships that\nare often correlated through time (e.g., 'woman, pay, money' followed by\n'woman, buy, coffee'). In this paper, we construct a Conditional Random Field\non a fully-connected spatio-temporal graph that exploits the statistical\ndependency between relational entities spatially and temporally. We introduce a\nnovel gated energy function parametrization that learns adaptive relations\nconditioned on visual observations. Our model optimization is computationally\nefficient, and its space computation complexity is significantly amortized\nthrough our proposed parameterization. Experimental results on benchmark video\ndatasets (ImageNet Video and Charades) demonstrate state-of-the-art performance\nacross three standard relationship reasoning tasks: Detection, Tagging, and\nRecognition.","url_abs":"http://arxiv.org/abs/1903.10547v2","url_pdf":"http://arxiv.org/pdf/1903.10547v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"video-relationship-reasoning-using-gated","repo_url":"https://github.com/yaohungt/GSTEG_CVPR_2019","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"model-optimization","task_name":"Model Optimization"},{"task_slug":"video-relationship","task_name":"Video Relationship"},{"task_slug":"video-visual-relation-detection","task_name":"Video Visual Relation Detection"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=1903.10547","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}