{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/videos-as-space-time-region-graphs","title":"Videos as Space-Time Region Graphs","arxiv_id":"1806.01810","date":"2018-06-05","proceeding":"ECCV 2018 9","authors":["Xiaolong Wang","Abhinav Gupta"],"abstract":"How do humans recognize the action \"opening a book\" ? We argue that there are\ntwo important cues: modeling temporal shape dynamics and modeling functional\nrelationships between humans and objects. In this paper, we propose to\nrepresent videos as space-time region graphs which capture these two important\ncues. Our graph nodes are defined by the object region proposals from different\nframes in a long range video. These nodes are connected by two types of\nrelations: (i) similarity relations capturing the long range dependencies\nbetween correlated objects and (ii) spatial-temporal relations capturing the\ninteractions between nearby objects. We perform reasoning on this graph\nrepresentation via Graph Convolutional Networks. We achieve state-of-the-art\nresults on both Charades and Something-Something datasets. Especially for\nCharades, we obtain a huge 4.4% gain when our model is applied in complex\nenvironments.","url_abs":"http://arxiv.org/abs/1806.01810v2","url_pdf":"http://arxiv.org/pdf/1806.01810v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"action-classification","task_name":"Action Classification"},{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"}],"methods":[{"method_slug":"graph-convolutional-networks","method_name":"Graph Convolutional Networks"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/action-classification-on-charades","task":"Action Classification","dataset":"Charades","model":"STRG","rank_in_archive_order":34,"of":49,"metrics":{"MAP":"39.7"},"uses_additional_data":true},{"leaderboard":"/sota/action-recognition-in-videos-on-something-1","task":"Action Recognition","dataset":"Something-Something V1","model":"NL I3D + GCN","rank_in_archive_order":68,"of":74,"metrics":{"Top 1 Accuracy":"46.1"},"uses_additional_data":false}],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=1806.01810","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}