{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/visual-interaction-networks","title":"Visual Interaction Networks","arxiv_id":"1706.01433","date":"2017-06-05","proceeding":null,"authors":["Nicholas Watters","Andrea Tacchetti","Theophane Weber","Razvan Pascanu","Peter Battaglia","Daniel Zoran"],"abstract":"From just a glance, humans can make rich predictions about the future state\nof a wide range of physical systems. On the other hand, modern approaches from\nengineering, robotics, and graphics are often restricted to narrow domains and\nrequire direct measurements of the underlying states. We introduce the Visual\nInteraction Network, a general-purpose model for learning the dynamics of a\nphysical system from raw visual observations. Our model consists of a\nperceptual front-end based on convolutional neural networks and a dynamics\npredictor based on interaction networks. Through joint training, the perceptual\nfront-end learns to parse a dynamic visual scene into a set of factored latent\nobject representations. The dynamics predictor learns to roll these states\nforward in time by computing their interactions and dynamics, producing a\npredicted physical trajectory of arbitrary length. We found that from just six\ninput video frames the Visual Interaction Network can generate accurate future\ntrajectories of hundreds of time steps on a wide range of physical systems. Our\nmodel can also be applied to scenes with invisible objects, inferring their\nfuture states from their effects on the visible objects, and can implicitly\ninfer the unknown mass of objects. Our results demonstrate that the perceptual\nmodule and the object-based dynamics predictor module can induce factored\nlatent representations that support accurate dynamical predictions. This work\nopens new opportunities for model-based decision-making and planning from raw\nsensory observations in complex physical environments.","url_abs":"http://arxiv.org/abs/1706.01433v1","url_pdf":"http://arxiv.org/pdf/1706.01433v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"visual-interaction-networks","repo_url":"https://github.com/clvrai/relation-network-tensorflow","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":null},{"paper_slug":"visual-interaction-networks","repo_url":"https://github.com/gitlimlab/Relation-Network-Tensorflow","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"unanswered"}},{"paper_slug":"visual-interaction-networks","repo_url":"https://github.com/jaesik817/visual-interaction-networks_tensorflow","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"decision-making","task_name":"Decision Making"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1706.01433","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}