{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/towards-good-practices-for-very-deep-two","title":"Towards Good Practices for Very Deep Two-Stream ConvNets","arxiv_id":"1507.02159","date":"2015-07-08","proceeding":null,"authors":["Limin Wang","Yuanjun Xiong","Zhe Wang","Yu Qiao"],"abstract":"Deep convolutional networks have achieved great success for object\nrecognition in still images. However, for action recognition in videos, the\nimprovement of deep convolutional networks is not so evident. We argue that\nthere are two reasons that could probably explain this result. First the\ncurrent network architectures (e.g. Two-stream ConvNets) are relatively shallow\ncompared with those very deep models in image domain (e.g. VGGNet, GoogLeNet),\nand therefore their modeling capacity is constrained by their depth. Second,\nprobably more importantly, the training dataset of action recognition is\nextremely small compared with the ImageNet dataset, and thus it will be easy to\nover-fit on the training dataset.\n  To address these issues, this report presents very deep two-stream ConvNets\nfor action recognition, by adapting recent very deep architectures into video\ndomain. However, this extension is not easy as the size of action recognition\nis quite small. We design several good practices for the training of very deep\ntwo-stream ConvNets, namely (i) pre-training for both spatial and temporal\nnets, (ii) smaller learning rates, (iii) more data augmentation techniques,\n(iv) high drop out ratio. Meanwhile, we extend the Caffe toolbox into Multi-GPU\nimplementation with high computational efficiency and low memory consumption.\nWe verify the performance of very deep two-stream ConvNets on the dataset of\nUCF101 and it achieves the recognition accuracy of $91.4\\%$.","url_abs":"http://arxiv.org/abs/1507.02159v1","url_pdf":"http://arxiv.org/pdf/1507.02159v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"towards-good-practices-for-very-deep-two","repo_url":"https://github.com/yjxiong/caffe","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":{"status":"unanswered"}},{"paper_slug":"towards-good-practices-for-very-deep-two","repo_url":"https://github.com/AbdalaDiasse/Video-classification-for-oil-quality-estimation","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"unanswered"}},{"paper_slug":"towards-good-practices-for-very-deep-two","repo_url":"https://github.com/bryanyzhu/two-stream-pytorch","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"unanswered"}},{"paper_slug":"towards-good-practices-for-very-deep-two","repo_url":"https://github.com/craftGBD/caffe-GBD","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null},{"paper_slug":"towards-good-practices-for-very-deep-two","repo_url":"https://github.com/gpspelle/Convert_channel","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"action-recognition-in-videos-2","task_name":"Action Recognition In Videos"},{"task_slug":"computational-efficiency","task_name":"Computational Efficiency"},{"task_slug":"data-augmentation","task_name":"Data Augmentation"},{"task_slug":null,"task_name":"GPU"},{"task_slug":"action-recognition","task_name":"Temporal Action Localization"},{"task_slug":"two","task_name":"Vocal Bursts Valence Prediction"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/action-recognition-in-videos-on-ucf101","task":"Action Recognition","dataset":"UCF101","model":"Very deep two-stream ConvNet","rank_in_archive_order":69,"of":91,"metrics":{"3-fold Accuracy":"91.4"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/1507.02159","atlas_url":"https://app.syntology.ai/?focus=1507.02159","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}