{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/action-search-spotting-actions-in-videos-and","title":"Action Search: Spotting Actions in Videos and Its Application to Temporal Action Localization","arxiv_id":"1706.04269","date":"2017-06-13","proceeding":"ECCV 2018 9","authors":["Humam Alwassel","Fabian Caba Heilbron","Bernard Ghanem"],"abstract":"State-of-the-art temporal action detectors inefficiently search the entire\nvideo for specific actions. Despite the encouraging progress these methods\nachieve, it is crucial to design automated approaches that only explore parts\nof the video which are the most relevant to the actions being searched for. To\naddress this need, we propose the new problem of action spotting in video,\nwhich we define as finding a specific action in a video while observing a small\nportion of that video. Inspired by the observation that humans are extremely\nefficient and accurate in spotting and finding action instances in video, we\npropose Action Search, a novel Recurrent Neural Network approach that mimics\nthe way humans spot actions. Moreover, to address the absence of data recording\nthe behavior of human annotators, we put forward the Human Searches dataset,\nwhich compiles the search sequences employed by human annotators spotting\nactions in the AVA and THUMOS14 datasets. We consider temporal action\nlocalization as an application of the action spotting problem. Experiments on\nthe THUMOS14 dataset reveal that our model is not only able to explore the\nvideo efficiently (observing on average 17.3% of the video) but it also\naccurately finds human activities with 30.8% mAP.","url_abs":"http://arxiv.org/abs/1706.04269v2","url_pdf":"http://arxiv.org/pdf/1706.04269v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"action-search-spotting-actions-in-videos-and","repo_url":"https://github.com/HumamAlwassel/action-search","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"action-localization","task_name":"Action Localization"},{"task_slug":"action-spotting","task_name":"Action Spotting"},{"task_slug":"action-recognition","task_name":"Temporal Action Localization"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}