{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/hardware-conditioned-policies-for-multi-robot","title":"Hardware Conditioned Policies for Multi-Robot Transfer Learning","arxiv_id":"1811.09864","date":"2018-11-24","proceeding":"NeurIPS 2018 12","authors":["Tao Chen","Adithyavairavan Murali","Abhinav Gupta"],"abstract":"Deep reinforcement learning could be used to learn dexterous robotic policies\nbut it is challenging to transfer them to new robots with vastly different\nhardware properties. It is also prohibitively expensive to learn a new policy\nfrom scratch for each robot hardware due to the high sample complexity of\nmodern state-of-the-art algorithms. We propose a novel approach called\n\\textit{Hardware Conditioned Policies} where we train a universal policy\nconditioned on a vector representation of robot hardware. We considered robots\nin simulation with varied dynamics, kinematic structure, kinematic lengths and\ndegrees-of-freedom. First, we use the kinematic structure directly as the\nhardware encoding and show great zero-shot transfer to completely novel robots\nnot seen during training. For robots with lower zero-shot success rate, we also\ndemonstrate that fine-tuning the policy network is significantly more\nsample-efficient than training a model from scratch. In tasks where knowing the\nagent dynamics is important for success, we learn an embedding for robot\nhardware and show that policies conditioned on the encoding of hardware tend to\ngeneralize and transfer well. The code and videos are available on the project\nwebpage: https://sites.google.com/view/robot-transfer-hcp.","url_abs":"http://arxiv.org/abs/1811.09864v2","url_pdf":"http://arxiv.org/pdf/1811.09864v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"hardware-conditioned-policies-for-multi-robot","repo_url":"https://github.com/taochenshh/hcp","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"deep-reinforcement-learning","task_name":"Deep Reinforcement Learning"},{"task_slug":"industrial-robots","task_name":"Industrial Robots"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"transfer-learning","task_name":"Transfer Learning"},{"task_slug":"transfer-reinforcement-learning","task_name":"Transfer Reinforcement Learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1811.09864","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}