{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/norml-no-reward-meta-learning","title":"NoRML: No-Reward Meta Learning","arxiv_id":"1903.01063","date":"2019-03-04","proceeding":null,"authors":["Yuxiang Yang","Ken Caluwaerts","Atil Iscen","Jie Tan","Chelsea Finn"],"abstract":"Efficiently adapting to new environments and changes in dynamics is critical\nfor agents to successfully operate in the real world. Reinforcement learning\n(RL) based approaches typically rely on external reward feedback for\nadaptation. However, in many scenarios this reward signal might not be readily\navailable for the target task, or the difference between the environments can\nbe implicit and only observable from the dynamics. To this end, we introduce a\nmethod that allows for self-adaptation of learned policies: No-Reward Meta\nLearning (NoRML). NoRML extends Model Agnostic Meta Learning (MAML) for RL and\nuses observable dynamics of the environment instead of an explicit reward\nfunction in MAML's finetune step. Our method has a more expressive update step\nthan MAML, while maintaining MAML's gradient based foundation. Additionally, in\norder to allow more targeted exploration, we implement an extension to MAML\nthat effectively disconnects the meta-policy parameters from the fine-tuned\npolicies' parameters. We first study our method on a number of synthetic\ncontrol problems and then validate our method on common benchmark environments,\nshowing that NoRML outperforms MAML when the dynamics change between tasks.","url_abs":"http://arxiv.org/abs/1903.01063v1","url_pdf":"http://arxiv.org/pdf/1903.01063v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"norml-no-reward-meta-learning","repo_url":"https://github.com/google-research/google-research","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"tf","reach":null}],"tasks":[{"task_slug":"meta-learning","task_name":"Meta-Learning"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"}],"methods":[{"method_slug":"maml","method_name":"MAML"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1903.01063","atlas_url":"https://app.syntology.ai/?focus=1903.01063","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}