{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/reward-estimation-for-variance-reduction-in","title":"Reward Estimation for Variance Reduction in Deep Reinforcement Learning","arxiv_id":"1805.03359","date":"2018-05-09","proceeding":null,"authors":["Joshua Romoff","Peter Henderson","Alexandre Piché","Vincent Francois-Lavet","Joelle Pineau"],"abstract":"Reinforcement Learning (RL) agents require the specification of a reward\nsignal for learning behaviours. However, introduction of corrupt or stochastic\nrewards can yield high variance in learning. Such corruption may be a direct\nresult of goal misspecification, randomness in the reward signal, or\ncorrelation of the reward with external factors that are not known to the\nagent. Corruption or stochasticity of the reward signal can be especially\nproblematic in robotics, where goal specification can be particularly difficult\nfor complex tasks. While many variance reduction techniques have been studied\nto improve the robustness of the RL process, handling such stochastic or\ncorrupted reward structures remains difficult. As an alternative for handling\nthis scenario in model-free RL methods, we suggest using an estimator for both\nrewards and value functions. We demonstrate that this improves performance\nunder corrupted stochastic rewards in both the tabular and non-linear function\napproximation settings for a variety of noise types and environments. The use\nof reward estimation is a robust and easy-to-implement improvement for handling\ncorrupted reward signals in model-free RL.","url_abs":"http://arxiv.org/abs/1805.03359v2","url_pdf":"http://arxiv.org/pdf/1805.03359v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"reward-estimation-for-variance-reduction-in","repo_url":"https://github.com/facebookresearch/reward-estimator-corl","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":{"status":"ok","spdx":"NOASSERTION"}}],"tasks":[{"task_slug":"deep-reinforcement-learning","task_name":"Deep Reinforcement Learning"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1805.03359","atlas_url":"https://app.syntology.ai/?focus=1805.03359","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}