{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/risk-sensitive-inverse-reinforcement-learning","title":"Risk-sensitive Inverse Reinforcement Learning via Semi- and Non-Parametric Methods","arxiv_id":"1711.10055","date":"2017-11-28","proceeding":null,"authors":["Sumeet Singh","Jonathan Lacotte","Anirudha Majumdar","Marco Pavone"],"abstract":"The literature on Inverse Reinforcement Learning (IRL) typically assumes that\nhumans take actions in order to minimize the expected value of a cost function,\ni.e., that humans are risk neutral. Yet, in practice, humans are often far from\nbeing risk neutral. To fill this gap, the objective of this paper is to devise\na framework for risk-sensitive IRL in order to explicitly account for a human's\nrisk sensitivity. To this end, we propose a flexible class of models based on\ncoherent risk measures, which allow us to capture an entire spectrum of risk\npreferences from risk-neutral to worst-case. We propose efficient\nnon-parametric algorithms based on linear programming and semi-parametric\nalgorithms based on maximum likelihood for inferring a human's underlying risk\nmeasure and cost function for a rich class of static and dynamic\ndecision-making settings. The resulting approach is demonstrated on a simulated\ndriving game with ten human participants. Our method is able to infer and mimic\na wide range of qualitatively different driving styles from highly risk-averse\nto risk-neutral in a data-efficient manner. Moreover, comparisons of the\nRisk-Sensitive (RS) IRL approach with a risk-neutral model show that the RS-IRL\nframework more accurately captures observed participant behavior both\nqualitatively and quantitatively, especially in scenarios where catastrophic\noutcomes such as collisions can occur.","url_abs":"http://arxiv.org/abs/1711.10055v2","url_pdf":"http://arxiv.org/pdf/1711.10055v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"risk-sensitive-inverse-reinforcement-learning","repo_url":"https://github.com/StanfordASL/RSIRL","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"decision-making","task_name":"Decision Making"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1711.10055","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}