{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/compressed-video-action-recognition","title":"Compressed Video Action Recognition","arxiv_id":"1712.00636","date":"2017-12-02","proceeding":"CVPR 2018 6","authors":["Chao-yuan Wu","Manzil Zaheer","Hexiang Hu","R. Manmatha","Alexander J. Smola","Philipp Krähenbühl"],"abstract":"Training robust deep video representations has proven to be much more\nchallenging than learning deep image representations. This is in part due to\nthe enormous size of raw video streams and the high temporal redundancy; the\ntrue and interesting signal is often drowned in too much irrelevant data.\nMotivated by that the superfluous information can be reduced by up to two\norders of magnitude by video compression (using H.264, HEVC, etc.), we propose\nto train a deep network directly on the compressed video.\n  This representation has a higher information density, and we found the\ntraining to be easier. In addition, the signals in a compressed video provide\nfree, albeit noisy, motion information. We propose novel techniques to use them\neffectively. Our approach is about 4.6 times faster than Res3D and 2.7 times\nfaster than ResNet-152. On the task of action recognition, our approach\noutperforms all the other methods on the UCF-101, HMDB-51, and Charades\ndataset.","url_abs":"http://arxiv.org/abs/1712.00636v2","url_pdf":"http://arxiv.org/pdf/1712.00636v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"compressed-video-action-recognition","repo_url":"https://github.com/chaoyuaw/pytorch-coviar","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"action-classification","task_name":"Action Classification"},{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"action-recognition","task_name":"Temporal Action Localization"},{"task_slug":"video-compression","task_name":"Video Compression"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/action-classification-on-charades","task":"Action Classification","dataset":"Charades","model":"CoViAR","rank_in_archive_order":46,"of":49,"metrics":{"MAP":"21.9"},"uses_additional_data":true}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1712.00636","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}