{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/guess-where-actor-supervision-for","title":"Guess Where? Actor-Supervision for Spatiotemporal Action Localization","arxiv_id":"1804.01824","date":"2018-04-05","proceeding":null,"authors":["Victor Escorcia","Cuong D. Dao","Mihir Jain","Bernard Ghanem","Cees Snoek"],"abstract":"This paper addresses the problem of spatiotemporal localization of actions in\nvideos. Compared to leading approaches, which all learn to localize based on\ncarefully annotated boxes on training video frames, we adhere to a\nweakly-supervised solution that only requires a video class label. We introduce\nan actor-supervised architecture that exploits the inherent compositionality of\nactions in terms of actor transformations, to localize actions. We make two\ncontributions. First, we propose actor proposals derived from a detector for\nhuman and non-human actors intended for images, which is linked over time by\nSiamese similarity matching to account for actor deformations. Second, we\npropose an actor-based attention mechanism that enables the localization of the\nactions from action class labels and actor proposals and is end-to-end\ntrainable. Experiments on three human and non-human action datasets show actor\nsupervision is state-of-the-art for weakly-supervised action localization and\nis even competitive to some fully-supervised alternatives.","url_abs":"http://arxiv.org/abs/1804.01824v1","url_pdf":"http://arxiv.org/pdf/1804.01824v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"guess-where-actor-supervision-for","repo_url":"https://github.com/escorciav/Text-to-Clip_Retrieval","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null},{"paper_slug":"guess-where-actor-supervision-for","repo_url":"https://github.com/escorciav/roi_pooling","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"action-localization","task_name":"Action Localization"},{"task_slug":"weakly-supervised-action-localization","task_name":"Weakly Supervised Action Localization"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1804.01824","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}