{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/untrimmednets-for-weakly-supervised-action","title":"UntrimmedNets for Weakly Supervised Action Recognition and Detection","arxiv_id":"1703.03329","date":"2017-03-09","proceeding":"CVPR 2017 7","authors":["Limin Wang","Yuanjun Xiong","Dahua Lin","Luc van Gool"],"abstract":"Current action recognition methods heavily rely on trimmed videos for model\ntraining. However, it is expensive and time-consuming to acquire a large-scale\ntrimmed video dataset. This paper presents a new weakly supervised\narchitecture, called UntrimmedNet, which is able to directly learn action\nrecognition models from untrimmed videos without the requirement of temporal\nannotations of action instances. Our UntrimmedNet couples two important\ncomponents, the classification module and the selection module, to learn the\naction models and reason about the temporal duration of action instances,\nrespectively. These two components are implemented with feed-forward networks,\nand UntrimmedNet is therefore an end-to-end trainable architecture. We exploit\nthe learned models for action recognition (WSR) and detection (WSD) on the\nuntrimmed video datasets of THUMOS14 and ActivityNet. Although our UntrimmedNet\nonly employs weak supervision, our method achieves performance superior or\ncomparable to that of those strongly supervised approaches on these two\ndatasets.","url_abs":"http://arxiv.org/abs/1703.03329v2","url_pdf":"http://arxiv.org/pdf/1703.03329v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"untrimmednets-for-weakly-supervised-action","repo_url":"https://github.com/wanglimin/UntrimmedNet","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok"}},{"paper_slug":"untrimmednets-for-weakly-supervised-action","repo_url":"https://github.com/zhengshou/AutoLoc","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"action-recognition","task_name":"Temporal Action Localization"},{"task_slug":"weakly-supervised-action-localization","task_name":"Weakly Supervised Action Localization"},{"task_slug":"weakly-supervised-action-recognition","task_name":"Weakly-Supervised Action Recognition"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/action-classification-on-activitynet-12","task":"Action Classification","dataset":"ActivityNet-1.2","model":"UntrimmedNets","rank_in_archive_order":3,"of":3,"metrics":{"mAP":"87.7"},"uses_additional_data":false},{"leaderboard":"/sota/action-classification-on-thumos14","task":"Action Classification","dataset":"THUMOS’14","model":"UntrimmedNets","rank_in_archive_order":3,"of":3,"metrics":{"mAP":"82.2"},"uses_additional_data":false},{"leaderboard":"/sota/weakly-supervised-action-localization-on","task":"Weakly Supervised Action Localization","dataset":"THUMOS 2014","model":"UntrimmedNets","rank_in_archive_order":29,"of":30,"metrics":{"mAP@0.1:0.7":"-","mAP@0.5":"13.7"},"uses_additional_data":false}],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=1703.03329","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}