{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/active-exploration-in-parameterized","title":"Active exploration in parameterized reinforcement learning","arxiv_id":"1610.01986","date":"2016-10-06","proceeding":null,"authors":["Mehdi Khamassi","Costas Tzafestas"],"abstract":"Online model-free reinforcement learning (RL) methods with continuous actions\nare playing a prominent role when dealing with real-world applications such as\nRobotics. However, when confronted to non-stationary environments, these\nmethods crucially rely on an exploration-exploitation trade-off which is rarely\ndynamically and automatically adjusted to changes in the environment. Here we\npropose an active exploration algorithm for RL in structured (parameterized)\ncontinuous action space. This framework deals with a set of discrete actions,\neach of which is parameterized with continuous variables. Discrete exploration\nis controlled through a Boltzmann softmax function with an inverse temperature\n$\\beta$ parameter. In parallel, a Gaussian exploration is applied to the\ncontinuous action parameters. We apply a meta-learning algorithm based on the\ncomparison between variations of short-term and long-term reward running\naverages to simultaneously tune $\\beta$ and the width of the Gaussian\ndistribution from which continuous action parameters are drawn. When applied to\na simple virtual human-robot interaction task, we show that this algorithm\noutperforms continuous parameterized RL both without active exploration and\nwith active exploration based on uncertainty variations measured by a\nKalman-Q-learning algorithm.","url_abs":"http://arxiv.org/abs/1610.01986v1","url_pdf":"http://arxiv.org/pdf/1610.01986v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"active-exploration-in-parameterized","repo_url":"https://github.com/MehdiKhamassi/SocialMetaLearning","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"meta-learning","task_name":"Meta-Learning"},{"task_slug":"q-learning","task_name":"Q-Learning"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[{"method_slug":"gaussian-process","method_name":"Gaussian Process"},{"method_slug":"q-learning","method_name":"Q-Learning"},{"method_slug":"softmax","method_name":"Softmax"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}