{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/preparing-for-the-unknown-learning-a","title":"Preparing for the Unknown: Learning a Universal Policy with Online System Identification","arxiv_id":"1702.02453","date":"2017-02-08","proceeding":null,"authors":["Wenhao Yu","Jie Tan","C. Karen Liu","Greg Turk"],"abstract":"We present a new method of learning control policies that successfully\noperate under unknown dynamic models. We create such policies by leveraging a\nlarge number of training examples that are generated using a physical\nsimulator. Our system is made of two components: a Universal Policy (UP) and a\nfunction for Online System Identification (OSI). We describe our control policy\nas universal because it is trained over a wide array of dynamic models. These\nvariations in the dynamic model may include differences in mass and inertia of\nthe robots' components, variable friction coefficients, or unknown mass of an\nobject to be manipulated. By training the Universal Policy with this variation,\nthe control policy is prepared for a wider array of possible conditions when\nexecuted in an unknown environment. The second part of our system uses the\nrecent state and action history of the system to predict the dynamics model\nparameters mu. The value of mu from the Online System Identification is then\nprovided as input to the control policy (along with the system state).\nTogether, UP-OSI is a robust control policy that can be used across a wide\nrange of dynamic models, and that is also responsive to sudden changes in the\nenvironment. We have evaluated the performance of this system on a variety of\ntasks, including the problem of cart-pole swing-up, the double inverted\npendulum, locomotion of a hopper, and block-throwing of a manipulator. UP-OSI\nis effective at these tasks across a wide range of dynamic models. Moreover,\nwhen tested with dynamic models outside of the training range, UP-OSI\noutperforms the Universal Policy alone, even when UP is given the actual value\nof the model dynamics. In addition to the benefits of creating more robust\ncontrollers, UP-OSI also holds out promise of narrowing the Reality Gap between\nsimulated and real physical systems.","url_abs":"http://arxiv.org/abs/1702.02453v3","url_pdf":"http://arxiv.org/pdf/1702.02453v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"preparing-for-the-unknown-learning-a","repo_url":"https://github.com/vincentyu68/policy_transfer","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok"}}],"tasks":[{"task_slug":"friction","task_name":"Friction"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1702.02453","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}