{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/do-less-and-achieve-more-training-cnns-for","title":"Do Less and Achieve More: Training CNNs for Action Recognition Utilizing Action Images from the Web","arxiv_id":"1512.07155","date":"2015-12-22","proceeding":null,"authors":["Shugao Ma","Sarah Adel Bargal","Jianming Zhang","Leonid Sigal","Stan Sclaroff"],"abstract":"Recently, attempts have been made to collect millions of videos to train CNN\nmodels for action recognition in videos. However, curating such large-scale\nvideo datasets requires immense human labor, and training CNNs on millions of\nvideos demands huge computational resources. In contrast, collecting action\nimages from the Web is much easier and training on images requires much less\ncomputation. In addition, labeled web images tend to contain discriminative\naction poses, which highlight discriminative portions of a video's temporal\nprogression. We explore the question of whether we can utilize web action\nimages to train better CNN models for action recognition in videos. We collect\n23.8K manually filtered images from the Web that depict the 101 actions in the\nUCF101 action video dataset. We show that by utilizing web action images along\nwith videos in training, significant performance boosts of CNN models can be\nachieved. We then investigate the scalability of the process by leveraging\ncrawled web images (unfiltered) for UCF101 and ActivityNet. We replace 16.2M\nvideo frames by 393K unfiltered images and get comparable performance.","url_abs":"http://arxiv.org/abs/1512.07155v1","url_pdf":"http://arxiv.org/pdf/1512.07155v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"action-recognition-in-videos-2","task_name":"Action Recognition In Videos"},{"task_slug":"action-recognition","task_name":"Temporal Action Localization"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/action-recognition-in-videos-on-activitynet","task":"Action Recognition","dataset":"ActivityNet","model":"VGG19 + 393K webcam images","rank_in_archive_order":14,"of":16,"metrics":{"mAP":"53.8"},"uses_additional_data":true},{"leaderboard":"/sota/action-recognition-in-videos-on-activitynet","task":"Action Recognition","dataset":"ActivityNet","model":"VGG19","rank_in_archive_order":16,"of":16,"metrics":{"mAP":"52.3"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1512.07155","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}