{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/learning-spatio-temporal-features-with-3d","title":"Learning Spatio-Temporal Features with 3D Residual Networks for Action Recognition","arxiv_id":"1708.07632","date":"2017-08-25","proceeding":null,"authors":["Kensho Hara","Hirokatsu Kataoka","Yutaka Satoh"],"abstract":"Convolutional neural networks with spatio-temporal 3D kernels (3D CNNs) have\nan ability to directly extract spatio-temporal features from videos for action\nrecognition. Although the 3D kernels tend to overfit because of a large number\nof their parameters, the 3D CNNs are greatly improved by using recent huge\nvideo databases. However, the architecture of 3D CNNs is relatively shallow\nagainst to the success of very deep neural networks in 2D-based CNNs, such as\nresidual networks (ResNets). In this paper, we propose a 3D CNNs based on\nResNets toward a better action representation. We describe the training\nprocedure of our 3D ResNets in details. We experimentally evaluate the 3D\nResNets on the ActivityNet and Kinetics datasets. The 3D ResNets trained on the\nKinetics did not suffer from overfitting despite the large number of parameters\nof the model, and achieved better performance than relatively shallow networks,\nsuch as C3D. Our code and pretrained models (e.g. Kinetics and ActivityNet) are\npublicly available at https://github.com/kenshohara/3D-ResNets.","url_abs":"http://arxiv.org/abs/1708.07632v1","url_pdf":"http://arxiv.org/pdf/1708.07632v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"learning-spatio-temporal-features-with-3d","repo_url":"https://github.com/kenshohara/3D-ResNets","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"torch","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"hand-gesture-recognition-1","task_name":"Hand-Gesture Recognition"},{"task_slug":"action-recognition","task_name":"Temporal Action Localization"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=1708.07632","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}