{"url":"/method/merl","slug":"merl","name":"MeRL","full_name":"Meta Reward Learning","full_name_withheld":false,"description_markdown":"**Meta Reward Learning (MeRL)** is a meta-learning method for the problem of learning from sparse and underspecified rewards. For example, an agent receives a complex input, such as a natural language instruction, and needs to generate a complex response, such as an action sequence, while only receiving binary success-failure feedback. The key insight of MeRL in dealing with underspecified rewards is that spurious trajectories and programs that achieve accidental success are detrimental to the agent's generalization performance. For example, an agent might be able to solve a specific instance of the maze problem above. However, if it learns to perform spurious actions during training, it is likely to fail when provided with unseen instructions. To mitigate this issue, MeRL optimizes a more refined auxiliary reward function, which can differentiate between accidental and purposeful success based on features of action trajectories. The auxiliary reward is optimized by maximizing the trained agent's performance on a hold-out validation set via meta learning.","description_state":"present","introduced_year":null,"introduced_by":{"title":null,"paper":null,"first_author":null,"n_authors":0,"url_abs":null,"archive_paper_url":null},"source":{"url":"https://arxiv.org/abs/1902.07198v4","title":"Learning to Generalize from Sparse and Underspecified Rewards","url_on_a_paper_host":true},"code_snippet_url":"https://github.com/google-research/google-research/tree/master/meta_reward_learning","code_snippet_url_on_a_code_host":true,"categories":[{"area":"General","area_id":"general","collection":"Meta-Learning Algorithms","url":"/methods/category/meta-learning-algorithms","pwc_aliases":[]}],"n_papers_tagged":10,"archive_num_papers":null,"papers_newest_first":[{"paper":null,"title":"Joint Antenna Position and Transmit Power Optimization for Pinching Antenna-Assisted ISAC Systems","date":"2025-03-17","arxiv_id":"2503.12872","n_code_links":0,"syntology":null},{"paper":"/paper/zero-shot-ecg-classification-with-multimodal","title":"Zero-Shot ECG Classification with Multimodal Learning and Test-time Clinical Knowledge Enhancement","date":"2024-03-11","arxiv_id":"2403.06659","n_code_links":2,"syntology":{"ran":8,"of":12,"unverified":4,"pointer_only":0}},{"paper":null,"title":"Jointly Learning Representations for Map Entities via Heterogeneous Graph Contrastive Learning","date":"2024-02-09","arxiv_id":"2402.06135","n_code_links":0,"syntology":null},{"paper":null,"title":"Hierarchical Attention Network for Action Segmentation","date":"2020-05-07","arxiv_id":"2005.03209","n_code_links":0,"syntology":null},{"paper":null,"title":"MERL: Multi-Head Reinforcement Learning","date":"2019-09-26","arxiv_id":"1909.11939","n_code_links":0,"syntology":null},{"paper":null,"title":"Coupled Generative Adversarial Network for Continuous Fine-grained Action Segmentation","date":"2019-09-20","arxiv_id":"1909.09283","n_code_links":0,"syntology":null},{"paper":null,"title":"Fine-grained Action Segmentation using the Semi-Supervised Action GAN","date":"2019-09-20","arxiv_id":"1909.09269","n_code_links":0,"syntology":null},{"paper":null,"title":"Evolutionary Reinforcement Learning for Sample-Efficient Multiagent Coordination","date":"2019-06-18","arxiv_id":"1906.07315","n_code_links":0,"syntology":null},{"paper":null,"title":"Follow the Attention: Combining Partial Pose and Object Motion for Fine-Grained Action Detection","date":"2019-05-11","arxiv_id":"1905.04430","n_code_links":0,"syntology":null},{"paper":"/paper/learning-to-generalize-from-sparse-and","title":"Learning to Generalize from Sparse and Underspecified Rewards","date":"2019-02-19","arxiv_id":"1902.07198","n_code_links":1,"syntology":null}],"papers_shown":10,"tasks":[{"task":"/task/action-segmentation","name":"Action Segmentation","papers":3},{"task":null,"name":"Generative Adversarial Network","papers":3},{"task":"/task/reinforcement-learning","name":"Reinforcement Learning","papers":2},{"task":"/task/reinforcement-learning-1","name":"Reinforcement Learning (RL)","papers":2},{"task":"/task/representation-learning","name":"Representation Learning","papers":2},{"task":"/task/segmentation","name":"Segmentation","papers":2},{"task":"/task/reinforcement-learning-2","name":"reinforcement-learning","papers":2},{"task":"/task/action-classification","name":"Action Classification","papers":1},{"task":"/task/action-detection","name":"Action Detection","papers":1},{"task":"/task/activity-detection","name":"Activity Detection","papers":1},{"task":"/task/activity-recognition","name":"Activity Recognition","papers":1},{"task":"/task/bayesian-optimization","name":"Bayesian Optimization","papers":1},{"task":"/task/clinical-knowledge","name":"Clinical Knowledge","papers":1},{"task":"/task/continuous-control","name":"Continuous Control","papers":1},{"task":"/task/contrastive-learning","name":"Contrastive Learning","papers":1},{"task":"/task/descriptive","name":"Descriptive","papers":1},{"task":"/task/diagnostic","name":"Diagnostic","papers":1},{"task":"/task/ecg-classification","name":"ECG Classification","papers":1},{"task":"/task/fine-grained-action-detection","name":"Fine-Grained Action Detection","papers":1},{"task":"/task/isac","name":"ISAC","papers":1}],"tasks_shown":20,"n_tasks":30,"usage_by_year":[{"year":"2019","papers":6},{"year":"2020","papers":1},{"year":"2024","papers":2},{"year":"2025","papers":1}],"row_source":"embedded","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/merl"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}