{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/joint-discovery-of-object-states-and","title":"Joint Discovery of Object States and Manipulation Actions","arxiv_id":"1702.02738","date":"2017-02-09","proceeding":"ICCV 2017 10","authors":["Jean-Baptiste Alayrac","Josev Sivic","Ivan Laptev","Simon Lacoste-Julien"],"abstract":"Many human activities involve object manipulations aiming to modify the\nobject state. Examples of common state changes include full/empty bottle,\nopen/closed door, and attached/detached car wheel. In this work, we seek to\nautomatically discover the states of objects and the associated manipulation\nactions. Given a set of videos for a particular task, we propose a joint model\nthat learns to identify object states and to localize state-modifying actions.\nOur model is formulated as a discriminative clustering cost with constraints.\nWe assume a consistent temporal order for the changes in object states and\nmanipulation actions, and introduce new optimization techniques to learn model\nparameters without additional supervision. We demonstrate successful discovery\nof seven manipulation actions and corresponding object states on a new dataset\nof videos depicting real-life object manipulations. We show that our joint\nformulation results in an improvement of object state discovery by action\nrecognition and vice versa.","url_abs":"http://arxiv.org/abs/1702.02738v3","url_pdf":"http://arxiv.org/pdf/1702.02738v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"joint-discovery-of-object-states-and","repo_url":"https://github.com/jalayrac/object-states-action","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"none","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"clustering","task_name":"Clustering"},{"task_slug":"object","task_name":"Object"},{"task_slug":"action-recognition","task_name":"Temporal Action Localization"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1702.02738","atlas_url":"https://app.syntology.ai/?focus=1702.02738","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}