{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/learning-spatiotemporal-features-with-3d","title":"Learning Spatiotemporal Features with 3D Convolutional Networks","arxiv_id":"1412.0767","date":"2014-12-02","proceeding":"ICCV 2015 12","authors":["Du Tran","Lubomir Bourdev","Rob Fergus","Lorenzo Torresani","Manohar Paluri"],"abstract":"We propose a simple, yet effective approach for spatiotemporal feature\nlearning using deep 3-dimensional convolutional networks (3D ConvNets) trained\non a large scale supervised video dataset. Our findings are three-fold: 1) 3D\nConvNets are more suitable for spatiotemporal feature learning compared to 2D\nConvNets; 2) A homogeneous architecture with small 3x3x3 convolution kernels in\nall layers is among the best performing architectures for 3D ConvNets; and 3)\nOur learned features, namely C3D (Convolutional 3D), with a simple linear\nclassifier outperform state-of-the-art methods on 4 different benchmarks and\nare comparable with current best methods on the other 2 benchmarks. In\naddition, the features are compact: achieving 52.8% accuracy on UCF101 dataset\nwith only 10 dimensions and also very efficient to compute due to the fast\ninference of ConvNets. Finally, they are conceptually very simple and easy to\ntrain and use.","url_abs":"http://arxiv.org/abs/1412.0767v4","url_pdf":"http://arxiv.org/pdf/1412.0767v4.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"learning-spatiotemporal-features-with-3d","repo_url":"https://github.com/facebookarchive/C3D","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"caffe2","reach":null},{"paper_slug":"learning-spatiotemporal-features-with-3d","repo_url":"https://github.com/AKASH2907/Content-based-Video-Recommendation","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok"}},{"paper_slug":"learning-spatiotemporal-features-with-3d","repo_url":"https://github.com/AKASH2907/Content-based-Video-Relevance-Prediction","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok"}},{"paper_slug":"learning-spatiotemporal-features-with-3d","repo_url":"https://github.com/MarkoLewis-Projects/Sign_language_detection","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"unanswered"}},{"paper_slug":"learning-spatiotemporal-features-with-3d","repo_url":"https://github.com/MekkaSiekka/C3D-UCF11-Tensorflow","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":null},{"paper_slug":"learning-spatiotemporal-features-with-3d","repo_url":"https://github.com/MichiganCOG/M-PACT","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"learning-spatiotemporal-features-with-3d","repo_url":"https://github.com/VEDANTGHODKE/Hand-Gesture-Recognition-Using-Neural-Networks","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null},{"paper_slug":"learning-spatiotemporal-features-with-3d","repo_url":"https://github.com/aim3-ruc/youmakeup_challenge2022","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}},{"paper_slug":"learning-spatiotemporal-features-with-3d","repo_url":"https://github.com/aj9011/Car-Speed-Prediction","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}},{"paper_slug":"learning-spatiotemporal-features-with-3d","repo_url":"https://github.com/ashu5711/Neural_Network_Hand_Gesture_Recognition","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null},{"paper_slug":"learning-spatiotemporal-features-with-3d","repo_url":"https://github.com/axon-research/c3d-keras","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"caffe2","reach":null},{"paper_slug":"learning-spatiotemporal-features-with-3d","repo_url":"https://github.com/coderSkyChen/Action_Recognition_Zoo","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"learning-spatiotemporal-features-with-3d","repo_url":"https://github.com/labs12/Action-Recgontion-","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"learning-spatiotemporal-features-with-3d","repo_url":"https://github.com/leftthomas/r2plus1d-c3d","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"learning-spatiotemporal-features-with-3d","repo_url":"https://github.com/mamtajha-ts/gesture-recognition","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null},{"paper_slug":"learning-spatiotemporal-features-with-3d","repo_url":"https://github.com/myaldiz/deep_violence_detection","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":null},{"paper_slug":"learning-spatiotemporal-features-with-3d","repo_url":"https://github.com/santhoshpkumar/Hand-gesture-recognition-using-neural-networks","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null},{"paper_slug":"learning-spatiotemporal-features-with-3d","repo_url":"https://github.com/waynshang/Gesture-Recognition-with-3DCNN","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null},{"paper_slug":"learning-spatiotemporal-features-with-3d","repo_url":"https://github.com/2023-MindSpore-1/ms-code-6/tree/main/C3D","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"mindspore","reach":null},{"paper_slug":"learning-spatiotemporal-features-with-3d","repo_url":"https://github.com/2024-MindSpore-1/Code5/tree/main/3dcnn","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"mindspore","reach":null},{"paper_slug":"learning-spatiotemporal-features-with-3d","repo_url":"https://github.com/2024-MindSpore-1/Code6/tree/main/3dcnn","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"mindspore","reach":null},{"paper_slug":"learning-spatiotemporal-features-with-3d","repo_url":"https://github.com/2024-MindSpore-1/Code6/tree/main/C3D","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"mindspore","reach":null},{"paper_slug":"learning-spatiotemporal-features-with-3d","repo_url":"https://github.com/HardyYoungX/C3D","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"mindspore","reach":null},{"paper_slug":"learning-spatiotemporal-features-with-3d","repo_url":"https://github.com/ZJUT-ERCISS/c3d_mindspore","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"mindspore","reach":null},{"paper_slug":"learning-spatiotemporal-features-with-3d","repo_url":"https://github.com/code-implementation1/Code2/tree/main/3dcnn","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"mindspore","reach":null},{"paper_slug":"learning-spatiotemporal-features-with-3d","repo_url":"https://github.com/mindspore-ai/models/tree/master/official/cv/C3D/src","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"mindspore","reach":null},{"paper_slug":"learning-spatiotemporal-features-with-3d","repo_url":"https://github.com/open-mmlab/mmaction2","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"pytorch","reach":null},{"paper_slug":"learning-spatiotemporal-features-with-3d","repo_url":"https://github.com/scouTT1/C3D","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"mindspore","reach":null},{"paper_slug":"learning-spatiotemporal-features-with-3d","repo_url":"https://github.com/xiuyu0000/vision/blob/main/mindvision/msvideo/models/c3d.py","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"mindspore","reach":null}],"tasks":[{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"action-recognition-in-videos-2","task_name":"Action Recognition In Videos"},{"task_slug":"dynamic-facial-expression-recognition","task_name":"Dynamic Facial Expression Recognition"}],"methods":[{"method_slug":"convolution","method_name":"Convolution"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/action-recognition-in-videos-on-hmdb-51","task":"Action Recognition","dataset":"HMDB-51","model":"C3D","rank_in_archive_order":75,"of":77,"metrics":{"Average accuracy of 3 splits":"51.6"},"uses_additional_data":false},{"leaderboard":"/sota/action-recognition-in-videos-on-sports-1m","task":"Action Recognition","dataset":"Sports-1M","model":"C3D","rank_in_archive_order":8,"of":9,"metrics":{"Clip Hit@1":"46.1","Video hit@1 ":"61.1","Video hit@5":"85.5"},"uses_additional_data":false},{"leaderboard":"/sota/action-recognition-in-videos-on-ucf101","task":"Action Recognition","dataset":"UCF101","model":"C3D","rank_in_archive_order":81,"of":91,"metrics":{"3-fold Accuracy":"82.3"},"uses_additional_data":false}],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=1412.0767","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}