{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/deep-reinforcement-learning-from-human","title":"Deep reinforcement learning from human preferences","arxiv_id":"1706.03741","date":"2017-06-12","proceeding":"NeurIPS 2017 12","authors":["Paul Christiano","Jan Leike","Tom B. Brown","Miljan Martic","Shane Legg","Dario Amodei"],"abstract":"For sophisticated reinforcement learning (RL) systems to interact usefully with real-world environments, we need to communicate complex goals to these systems. In this work, we explore goals defined in terms of (non-expert) human preferences between pairs of trajectory segments. We show that this approach can effectively solve complex RL tasks without access to the reward function, including Atari games and simulated robot locomotion, while providing feedback on less than one percent of our agent's interactions with the environment. This reduces the cost of human oversight far enough that it can be practically applied to state-of-the-art RL systems. To demonstrate the flexibility of our approach, we show that we can successfully train complex novel behaviors with about an hour of human time. These behaviors and environments are considerably more complex than any that have been previously learned from human feedback.","url_abs":"https://arxiv.org/abs/1706.03741v4","url_pdf":"https://arxiv.org/pdf/1706.03741v4.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"deep-reinforcement-learning-from-human","repo_url":"https://github.com/JulienDesvergnes/human-reinforcement-learning","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok"}},{"paper_slug":"deep-reinforcement-learning-from-human","repo_url":"https://github.com/RESQUELAB/RL-Teacher-UIAdaptation","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null},{"paper_slug":"deep-reinforcement-learning-from-human","repo_url":"https://github.com/ZachisGit/LearningFromHumanPreferences","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"unanswered"}},{"paper_slug":"deep-reinforcement-learning-from-human","repo_url":"https://github.com/kaichiuwong/rlhps","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null},{"paper_slug":"deep-reinforcement-learning-from-human","repo_url":"https://github.com/mrahtz/learning-from-human-preferences","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"unanswered"}},{"paper_slug":"deep-reinforcement-learning-from-human","repo_url":"https://github.com/vcharvet/project-rl","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"unanswered"}},{"paper_slug":"deep-reinforcement-learning-from-human","repo_url":"https://github.com/Kunal-Kumar-Sahoo/Deep-Reinforcement-Learning-with-Human-Preferences","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"none","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"atari-games","task_name":"Atari Games"},{"task_slug":"deep-reinforcement-learning","task_name":"Deep Reinforcement Learning"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1706.03741","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"1706.03741"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/vcharvet/project-rl","reach":{"status":"unanswered"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/RESQUELAB/RL-Teacher-UIAdaptation","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/kaichiuwong/rlhps","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/Kunal-Kumar-Sahoo/Deep-Reinforcement-Learning-with-Human-Preferences","reach":{"status":"unanswered"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/mrahtz/learning-from-human-preferences","reach":{"status":"unanswered"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/JulienDesvergnes/human-reinforcement-learning","reach":{"status":"ok"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/ZachisGit/LearningFromHumanPreferences","reach":{"status":"unanswered"}}],"summary":{"ran_draft_wrong":2,"unverified":1},"by_repo_kind":{"listed":{"samples":3,"ran":2,"repositories":2}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"768a7ed2c447a0e3","entry":"createConfigJSON","repo":"RESQUELAB/RL-Teacher-UIAdaptation","repo_kind":"listed","path":"rl_teacher/teach.py","file_url":"https://github.com/RESQUELAB/RL-Teacher-UIAdaptation/blob/HEAD/rl_teacher/teach.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"768a7ed2c447a0e3"}},{"code_sha256_prefix":"69003888fa70c03c","entry":"createUIDesign","repo":"RESQUELAB/RL-Teacher-UIAdaptation","repo_kind":"listed","path":"rl_teacher/teach.py","file_url":"https://github.com/RESQUELAB/RL-Teacher-UIAdaptation/blob/HEAD/rl_teacher/teach.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"69003888fa70c03c"}},{"code_sha256_prefix":"fa9273661876d886","entry":"wrap_train","repo":"kaichiuwong/rlhps","repo_kind":"listed","path":"agents/pposgd-mpi/pposgd_mpi/run_atari.py","file_url":"https://github.com/kaichiuwong/rlhps/blob/HEAD/agents/pposgd-mpi/pposgd_mpi/run_atari.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"fa9273661876d886"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}