{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/where-do-you-think-youre-going-inferring","title":"Where Do You Think You're Going?: Inferring Beliefs about Dynamics from Behavior","arxiv_id":"1805.08010","date":"2018-05-21","proceeding":"NeurIPS 2018 12","authors":["Siddharth Reddy","Anca D. Dragan","Sergey Levine"],"abstract":"Inferring intent from observed behavior has been studied extensively within\nthe frameworks of Bayesian inverse planning and inverse reinforcement learning.\nThese methods infer a goal or reward function that best explains the actions of\nthe observed agent, typically a human demonstrator. Another agent can use this\ninferred intent to predict, imitate, or assist the human user. However, a\ncentral assumption in inverse reinforcement learning is that the demonstrator\nis close to optimal. While models of suboptimal behavior exist, they typically\nassume that suboptimal actions are the result of some type of random noise or a\nknown cognitive bias, like temporal inconsistency. In this paper, we take an\nalternative approach, and model suboptimal behavior as the result of internal\nmodel misspecification: the reason that user actions might deviate from\nnear-optimal actions is that the user has an incorrect set of beliefs about the\nrules -- the dynamics -- governing how actions affect the environment. Our\ninsight is that while demonstrated actions may be suboptimal in the real world,\nthey may actually be near-optimal with respect to the user's internal model of\nthe dynamics. By estimating these internal beliefs from observed behavior, we\narrive at a new method for inferring intent. We demonstrate in simulation and\nin a user study with 12 participants that this approach enables us to more\naccurately model human intent, and can be used in a variety of applications,\nincluding offering assistance in a shared autonomy framework and inferring\nhuman preferences.","url_abs":"http://arxiv.org/abs/1805.08010v4","url_pdf":"http://arxiv.org/pdf/1805.08010v4.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"where-do-you-think-youre-going-inferring","repo_url":"https://github.com/rddy/isql","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":null}],"tasks":[{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1805.08010","atlas_url":"https://app.syntology.ai/?focus=1805.08010","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"1805.08010"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-25T09:33:49+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/rddy/isql","reach":null}],"summary":{"ran_fixture":1},"by_repo_kind":{"official":{"samples":1,"ran":1,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"750266fcc0999cc7","entry":"heuristic","repo":"rddy/isql","repo_kind":"official","path":"lunar_lander.py","file_url":"https://github.com/rddy/isql/blob/HEAD/lunar_lander.py","link_basis":"first_harvest_node","language":"python","status":"ran_fixture","verification_level":1,"contract_check":"RAISES","metamorphic_tier":"invariant","behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"750266fcc0999cc7"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}