{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/actionflownet-learning-motion-representation","title":"ActionFlowNet: Learning Motion Representation for Action Recognition","arxiv_id":"1612.03052","date":"2016-12-09","proceeding":null,"authors":["Joe Yue-Hei Ng","Jonghyun Choi","Jan Neumann","Larry S. Davis"],"abstract":"Even with the recent advances in convolutional neural networks (CNN) in\nvarious visual recognition tasks, the state-of-the-art action recognition\nsystem still relies on hand crafted motion feature such as optical flow to\nachieve the best performance. We propose a multitask learning model\nActionFlowNet to train a single stream network directly from raw pixels to\njointly estimate optical flow while recognizing actions with convolutional\nneural networks, capturing both appearance and motion in a single model. We\nadditionally provide insights to how the quality of the learned optical flow\naffects the action recognition. Our model significantly improves action\nrecognition accuracy by a large margin 31% compared to state-of-the-art\nCNN-based action recognition models trained without external large scale data\nand additional optical flow input. Without pretraining on large external\nlabeled datasets, our model, by well exploiting the motion information,\nachieves competitive recognition accuracy to the models trained with large\nlabeled datasets such as ImageNet and Sport-1M.","url_abs":"http://arxiv.org/abs/1612.03052v3","url_pdf":"http://arxiv.org/pdf/1612.03052v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"optical-flow-estimation","task_name":"Optical Flow Estimation"},{"task_slug":"action-recognition","task_name":"Temporal Action Localization"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/action-recognition-in-videos-on-hmdb-51","task":"Action Recognition","dataset":"HMDB-51","model":"ActionFlowNet","rank_in_archive_order":69,"of":77,"metrics":{"Average accuracy of 3 splits":"56.4"},"uses_additional_data":false},{"leaderboard":"/sota/action-recognition-in-videos-on-ucf101","task":"Action Recognition","dataset":"UCF101","model":"ActionFlowNet","rank_in_archive_order":80,"of":91,"metrics":{"3-fold Accuracy":"83.9"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1612.03052","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}