{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/actions-transformations","title":"Actions ~ Transformations","arxiv_id":"1512.00795","date":"2015-12-02","proceeding":"CVPR 2016 6","authors":["Xiaolong Wang","Ali Farhadi","Abhinav Gupta"],"abstract":"What defines an action like \"kicking ball\"? We argue that the true meaning of\nan action lies in the change or transformation an action brings to the\nenvironment. In this paper, we propose a novel representation for actions by\nmodeling an action as a transformation which changes the state of the\nenvironment before the action happens (precondition) to the state after the\naction (effect). Motivated by recent advancements of video representation using\ndeep learning, we design a Siamese network which models the action as a\ntransformation on a high-level feature space. We show that our model gives\nimprovements on standard action recognition datasets including UCF101 and\nHMDB51. More importantly, our approach is able to generalize beyond learned\naction categories and shows significant performance improvement on\ncross-category generalization on our new ACT dataset.","url_abs":"http://arxiv.org/abs/1512.00795v2","url_pdf":"http://arxiv.org/pdf/1512.00795v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"actions-transformations","repo_url":"https://github.com/andresespinosapc/actions-transformations-pytorch","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"action-recognition","task_name":"Temporal Action Localization"}],"methods":[{"method_slug":"siamese-network","method_name":"Siamese Network"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1512.00795","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}