{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/end-to-end-learning-of-action-detection-from","title":"End-to-end Learning of Action Detection from Frame Glimpses in Videos","arxiv_id":"1511.06984","date":"2015-11-22","proceeding":"CVPR 2016 6","authors":["Serena Yeung","Olga Russakovsky","Greg Mori","Li Fei-Fei"],"abstract":"In this work we introduce a fully end-to-end approach for action detection in\nvideos that learns to directly predict the temporal bounds of actions. Our\nintuition is that the process of detecting actions is naturally one of\nobservation and refinement: observing moments in video, and refining hypotheses\nabout when an action is occurring. Based on this insight, we formulate our\nmodel as a recurrent neural network-based agent that interacts with a video\nover time. The agent observes video frames and decides both where to look next\nand when to emit a prediction. Since backpropagation is not adequate in this\nnon-differentiable setting, we use REINFORCE to learn the agent's decision\npolicy. Our model achieves state-of-the-art results on the THUMOS'14 and\nActivityNet datasets while observing only a fraction (2% or less) of the video\nframes.","url_abs":"http://arxiv.org/abs/1511.06984v2","url_pdf":"http://arxiv.org/pdf/1511.06984v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"end-to-end-learning-of-action-detection-from","repo_url":"https://github.com/syyeung/frameglimpses","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"torch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"action-detection","task_name":"Action Detection"},{"task_slug":"action-recognition","task_name":"Temporal Action Localization"}],"methods":[{"method_slug":"reinforce","method_name":"REINFORCE"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/action-recognition-in-videos-on-thumos14","task":"Action Recognition","dataset":"THUMOS’14","model":"Yeung et. al.","rank_in_archive_order":10,"of":10,"metrics":{"mAP@0.1":"48.9","mAP@0.2":"44.0","mAP@0.3":"36.0","mAP@0.4":"26.4","mAP@0.5":"17.1"},"uses_additional_data":false},{"leaderboard":"/sota/temporal-action-localization-on-thumos14","task":"Temporal Action Localization","dataset":"THUMOS’14","model":"Yeung et al.","rank_in_archive_order":37,"of":42,"metrics":{"mAP IOU@0.1":"48.9","mAP IOU@0.2":"44.0","mAP IOU@0.3":"36.0","mAP IOU@0.4":"26.4","mAP IOU@0.5":"17.1"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/1511.06984","atlas_url":"https://app.syntology.ai/?focus=1511.06984","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}