{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/shapestacks-learning-vision-based-physical","title":"ShapeStacks: Learning Vision-Based Physical Intuition for Generalised Object Stacking","arxiv_id":"1804.08018","date":"2018-04-21","proceeding":"ECCV 2018 9","authors":["Oliver Groth","Fabian B. Fuchs","Ingmar Posner","Andrea Vedaldi"],"abstract":"Physical intuition is pivotal for intelligent agents to perform complex\ntasks. In this paper we investigate the passive acquisition of an intuitive\nunderstanding of physical principles as well as the active utilisation of this\nintuition in the context of generalised object stacking. To this end, we\nprovide: a simulation-based dataset featuring 20,000 stack configurations\ncomposed of a variety of elementary geometric primitives richly annotated\nregarding semantics and structural stability. We train visual classifiers for\nbinary stability prediction on the ShapeStacks data and scrutinise their\nlearned physical intuition. Due to the richness of the training data our\napproach also generalises favourably to real-world scenarios achieving\nstate-of-the-art stability prediction on a publicly available benchmark of\nblock towers. We then leverage the physical intuition learned by our model to\nactively construct stable stacks and observe the emergence of an intuitive\nnotion of stackability - an inherent object affordance - induced by the active\nstacking task. Our approach performs well even in challenging conditions where\nit considerably exceeds the stack height observed during training or in cases\nwhere initially unstable structures must be stabilised via counterbalancing.","url_abs":"http://arxiv.org/abs/1804.08018v2","url_pdf":"http://arxiv.org/pdf/1804.08018v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"shapestacks-learning-vision-based-physical","repo_url":"https://github.com/ogroth/shapestacks","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"tf","reach":{"status":"ok","spdx":"GPL-3.0"}}],"tasks":[{"task_slug":"physical-intuition","task_name":"Physical Intuition"}],"methods":[],"datasets_introduced":[{"slug":"shapestacks","name":"ShapeStacks","full_name":""}],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1804.08018","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}