{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/denseimage-network-video-spatial-temporal","title":"DenseImage Network: Video Spatial-Temporal Evolution Encoding and Understanding","arxiv_id":"1805.07550","date":"2018-05-19","proceeding":null,"authors":["Xiaokai Chen","Ke Gao"],"abstract":"Many of the leading approaches for video understanding are data-hungry and\ntime-consuming, failing to capture the gist of spatial-temporal evolution in an\nefficient manner. The latest research shows that CNN network can reason about\nstatic relation of entities in images. To further exploit its capacity in\ndynamic evolution reasoning, we introduce a novel network module called\nDenseImage Network(DIN) with two main contributions. 1) A novel compact\nrepresentation of video which distills its significant spatial-temporal\nevolution into a matrix called DenseImage, primed for efficient video encoding.\n2) A simple yet powerful learning strategy based on DenseImage and a\ntemporal-order-preserving CNN network is proposed for video understanding,\nwhich contains a local temporal correlation constraint capturing temporal\nevolution at multiple time scales with different filter widths. Extensive\nexperiments on two recent challenging benchmarks demonstrate that our\nDenseImage Network can accurately capture the common spatial-temporal evolution\nbetween similar actions, even with enormous visual variations or different time\nscales. Moreover, we obtain the state-of-the-art results in action and gesture\nrecognition with much less time-and-memory cost, indicating its immense\npotential in video representing and understanding.","url_abs":"http://arxiv.org/abs/1805.07550v1","url_pdf":"http://arxiv.org/pdf/1805.07550v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"action-recognition-in-videos-2","task_name":"Action Recognition In Videos"},{"task_slug":"gesture-recognition","task_name":"Gesture Recognition"},{"task_slug":"video-understanding","task_name":"Video Understanding"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/action-recognition-in-videos-on-jester-1","task":"Action Recognition In Videos","dataset":"Jester (Gesture Recognition)","model":"DIN","rank_in_archive_order":5,"of":9,"metrics":{"Val":"95.31"},"uses_additional_data":false},{"leaderboard":"/sota/action-recognition-in-videos-on-something-3","task":"Action Recognition In Videos","dataset":"Something-Something V2","model":"DIN","rank_in_archive_order":4,"of":4,"metrics":{"Top-1 Accuracy":"34.11"},"uses_additional_data":false}],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}