{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/unsupervised-intuitive-physics-from-visual","title":"Unsupervised Intuitive Physics from Visual Observations","arxiv_id":"1805.05086","date":"2018-05-14","proceeding":null,"authors":["Sebastien Ehrhardt","Aron Monszpart","Niloy Mitra","Andrea Vedaldi"],"abstract":"While learning models of intuitive physics is an increasingly active area of\nresearch, current approaches still fall short of natural intelligences in one\nimportant regard: they require external supervision, such as explicit access to\nphysical states, at training and sometimes even at test times. Some authors\nhave relaxed such requirements by supplementing the model with an handcrafted\nphysical simulator. Still, the resulting methods are unable to automatically\nlearn new complex environments and to understand physical interactions within\nthem. In this work, we demonstrated for the first time learning such predictors\ndirectly from raw visual observations and without relying on simulators. We do\nso in two steps: first, we learn to track mechanically-salient objects in\nvideos using causality and equivariance, two unsupervised learning principles\nthat do not require auto-encoding. Second, we demonstrate that the extracted\npositions are sufficient to successfully train visual motion predictors that\ncan take the underlying environment into account. We validate our predictors on\nsynthetic datasets; then, we introduce a new dataset, ROLL4REAL, consisting of\nreal objects rolling on complex terrains (pool table, elliptical bowl, and\nrandom height-field). We show that in all such cases it is possible to learn\nreliable extrapolators of the object trajectories from raw videos alone,\nwithout any form of external supervision and with no more prior knowledge than\nthe choice of a convolutional neural network architecture.","url_abs":"http://arxiv.org/abs/1805.05086v2","url_pdf":"http://arxiv.org/pdf/1805.05086v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[],"methods":[],"datasets_introduced":[{"slug":"roll4real","name":"Roll4Real","full_name":""}],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1805.05086","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}