{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/action-recognition-with-dynamic-image","title":"Action Recognition with Dynamic Image Networks","arxiv_id":"1612.00738","date":"2016-12-02","proceeding":null,"authors":["Hakan Bilen","Basura Fernando","Efstratios Gavves","Andrea Vedaldi"],"abstract":"We introduce the concept of \"dynamic image\", a novel compact representation\nof videos useful for video analysis, particularly in combination with\nconvolutional neural networks (CNNs). A dynamic image encodes temporal data\nsuch as RGB or optical flow videos by using the concept of `rank pooling'. The\nidea is to learn a ranking machine that captures the temporal evolution of the\ndata and to use the parameters of the latter as a representation. When a linear\nranking machine is used, the resulting representation is in the form of an\nimage, which we call dynamic because it summarizes the video dynamics in\naddition of appearance. This is a powerful idea because it allows to convert\nany video to an image so that existing CNN models pre-trained for the analysis\nof still images can be immediately extended to videos. We also present an\nefficient and effective approximate rank pooling operator, accelerating\nstandard rank pooling algorithms by orders of magnitude, and formulate that as\na CNN layer. This new layer allows generalizing dynamic images to dynamic\nfeature maps. We demonstrate the power of the new representations on standard\nbenchmarks in action recognition achieving state-of-the-art performance.","url_abs":"http://arxiv.org/abs/1612.00738v2","url_pdf":"http://arxiv.org/pdf/1612.00738v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"action-recognition-with-dynamic-image","repo_url":"https://github.com/hbilen/dynamic-image-nets","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null},{"paper_slug":"action-recognition-with-dynamic-image","repo_url":"https://github.com/Dartum08/Video-Anomaly-Detection","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":null},{"paper_slug":"action-recognition-with-dynamic-image","repo_url":"https://github.com/ayushraj7/Video-Anomaly-Detection","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":null}],"tasks":[{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"optical-flow-estimation","task_name":"Optical Flow Estimation"},{"task_slug":"action-recognition","task_name":"Temporal Action Localization"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}