{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/real-time-action-recognition-with-enhanced","title":"Real-time Action Recognition with Enhanced Motion Vector CNNs","arxiv_id":"1604.07669","date":"2016-04-26","proceeding":"CVPR 2016 6","authors":["Bowen Zhang","Li-Min Wang","Zhe Wang","Yu Qiao","Hanli Wang"],"abstract":"The deep two-stream architecture exhibited excellent performance on video\nbased action recognition. The most computationally expensive step in this\napproach comes from the calculation of optical flow which prevents it to be\nreal-time. This paper accelerates this architecture by replacing optical flow\nwith motion vector which can be obtained directly from compressed videos\nwithout extra calculation. However, motion vector lacks fine structures, and\ncontains noisy and inaccurate motion patterns, leading to the evident\ndegradation of recognition performance. Our key insight for relieving this\nproblem is that optical flow and motion vector are inherent correlated.\nTransferring the knowledge learned with optical flow CNN to motion vector CNN\ncan significantly boost the performance of the latter. Specifically, we\nintroduce three strategies for this, initialization transfer, supervision\ntransfer and their combination. Experimental results show that our method\nachieves comparable recognition performance to the state-of-the-art, while our\nmethod can process 390.7 frames per second, which is 27 times faster than the\noriginal two-stream method.","url_abs":"http://arxiv.org/abs/1604.07669v1","url_pdf":"http://arxiv.org/pdf/1604.07669v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"real-time-action-recognition-with-enhanced","repo_url":"https://github.com/yjxiong/caffe","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"optical-flow-estimation","task_name":"Optical Flow Estimation"},{"task_slug":"action-recognition","task_name":"Temporal Action Localization"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/action-recognition-in-videos-on-ucf101","task":"Action Recognition","dataset":"UCF101","model":"MV-CNN","rank_in_archive_order":76,"of":91,"metrics":{"3-fold Accuracy":"86.4"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1604.07669","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}