{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/action-recognition-for-depth-video-using","title":"Action Recognition for Depth Video using Multi-view Dynamic Images","arxiv_id":"1806.11269","date":"2018-06-29","proceeding":null,"authors":["Yang Xiao","Jun Chen","Yancheng Wang","Zhiguo Cao","Joey Tianyi Zhou","Xiang Bai"],"abstract":"Dynamic imaging is a recently proposed action description paradigm for\nsimultaneously capturing motion and temporal evolution information,\nparticularly in the context of deep convolutional neural networks (CNNs).\nCompared with optical flow for motion characterization, dynamic imaging\nexhibits superior efficiency and compactness. Inspired by the success of\ndynamic imaging in RGB video, this study extends it to the depth domain. To\nbetter exploit three-dimensional (3D) characteristics, multi-view dynamic\nimages are proposed. In particular, the raw depth video is densely projected\nwith respect to different virtual imaging viewpoints by rotating the virtual\ncamera within the 3D space. Subsequently, dynamic images are extracted from the\nobtained multi-view depth videos and multi-view dynamic images are thus\nconstructed from these images. Accordingly, more view-tolerant visual cues can\nbe involved. A novel CNN model is then proposed to perform feature learning on\nmulti-view dynamic images. Particularly, the dynamic images from different\nviews share the same convolutional layers but correspond to different fully\nconnected layers. This is aimed at enhancing the tuning effectiveness on\nshallow convolutional layers by alleviating the gradient vanishing problem.\nMoreover, as the spatial occurrence variation of the actions may impair the\nCNN, an action proposal approach is also put forth. In experiments, the\nproposed approach can achieve state-of-the-art performance on three challenging\ndatasets.","url_abs":"http://arxiv.org/abs/1806.11269v3","url_pdf":"http://arxiv.org/pdf/1806.11269v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"action-recognition-for-depth-video-using","repo_url":"https://github.com/3huo/MVDI","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":{"status":"ok"}}],"tasks":[{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"optical-flow-estimation","task_name":"Optical Flow Estimation"},{"task_slug":"action-recognition","task_name":"Temporal Action Localization"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1806.11269","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}