{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/learning-optimal-advantage-from-preferences","title":"Learning Optimal Advantage from Preferences and Mistaking it for Reward","arxiv_id":"2310.02456","date":"2023-10-03","proceeding":null,"authors":["W. Bradley Knox","Stephane Hatgis-Kessell","Sigurdur Orn Adalgeirsson","Serena Booth","Anca Dragan","Peter Stone","Scott Niekum"],"abstract":"We consider algorithms for learning reward functions from human preferences over pairs of trajectory segments, as used in reinforcement learning from human feedback (RLHF). Most recent work assumes that human preferences are generated based only upon the reward accrued within those segments, or their partial return. Recent work casts doubt on the validity of this assumption, proposing an alternative preference model based upon regret. We investigate the consequences of assuming preferences are based upon partial return when they actually arise from regret. We argue that the learned function is an approximation of the optimal advantage function, $\\hat{A^*_r}$, not a reward function. We find that if a specific pitfall is addressed, this incorrect assumption is not particularly harmful, resulting in a highly shaped reward function. Nonetheless, this incorrect usage of $\\hat{A^*_r}$ is less desirable than the appropriate and simpler approach of greedy maximization of $\\hat{A^*_r}$. From the perspective of the regret preference model, we also provide a clearer interpretation of fine tuning contemporary large language models with RLHF. This paper overall provides insight regarding why learning under the partial return preference model tends to work so well in practice, despite it conforming poorly to how humans give preferences.","url_abs":"https://arxiv.org/abs/2310.02456v1","url_pdf":"https://arxiv.org/pdf/2310.02456v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"learning-optimal-advantage-from-preferences","repo_url":"https://github.com/Stephanehk/Learning-OA-From-Prefs","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2310.02456","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2310.02456"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/Stephanehk/Learning-OA-From-Prefs","reach":{"status":"ok"}}],"summary":{"ran":11,"unverified":1},"by_repo_kind":{"official":{"samples":12,"ran":11,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":12,"samples":[{"code_sha256_prefix":"f6e5bcc0d13ef9a6","entry":"build_pi_from_nn_feats","repo":"Stephanehk/Learning-OA-From-Prefs","repo_kind":"official","path":"learn_advantage/algorithms/rl_algos.py","file_url":"https://github.com/Stephanehk/Learning-OA-From-Prefs/blob/HEAD/learn_advantage/algorithms/rl_algos.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"f6e5bcc0d13ef9a6"}},{"code_sha256_prefix":"7739f17cbdd84716","entry":"build_reward_from_nn_feats","repo":"Stephanehk/Learning-OA-From-Prefs","repo_kind":"official","path":"learn_advantage/algorithms/rl_algos.py","file_url":"https://github.com/Stephanehk/Learning-OA-From-Prefs/blob/HEAD/learn_advantage/algorithms/rl_algos.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"7739f17cbdd84716"}},{"code_sha256_prefix":"1d126d83c8e0f620","entry":"find_end_state","repo":"Stephanehk/Learning-OA-From-Prefs","repo_kind":"official","path":"learn_advantage/env/generate_random_mdp.py","file_url":"https://github.com/Stephanehk/Learning-OA-From-Prefs/blob/HEAD/learn_advantage/env/generate_random_mdp.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"1d126d83c8e0f620"}},{"code_sha256_prefix":"e3448e1f08ace362","entry":"get_loops","repo":"Stephanehk/Learning-OA-From-Prefs","repo_kind":"official","path":"learn_advantage/analysis/generate_fig_5.py","file_url":"https://github.com/Stephanehk/Learning-OA-From-Prefs/blob/HEAD/learn_advantage/analysis/generate_fig_5.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"e3448e1f08ace362"}},{"code_sha256_prefix":"45d3e57f998c2581","entry":"get_random_reward_vector","repo":"Stephanehk/Learning-OA-From-Prefs","repo_kind":"official","path":"learn_advantage/env/generate_random_policies.py","file_url":"https://github.com/Stephanehk/Learning-OA-From-Prefs/blob/HEAD/learn_advantage/env/generate_random_policies.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"45d3e57f998c2581"}},{"code_sha256_prefix":"7830dff13490ae72","entry":"get_reward_vec","repo":"Stephanehk/Learning-OA-From-Prefs","repo_kind":"official","path":"learn_advantage/algorithms/advantage_learning.py","file_url":"https://github.com/Stephanehk/Learning-OA-From-Prefs/blob/HEAD/learn_advantage/algorithms/advantage_learning.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"7830dff13490ae72"}},{"code_sha256_prefix":"0428bfbecb501b40","entry":"is_arr_in_list","repo":"Stephanehk/Learning-OA-From-Prefs","repo_kind":"official","path":"learn_advantage/env/generate_random_policies.py","file_url":"https://github.com/Stephanehk/Learning-OA-From-Prefs/blob/HEAD/learn_advantage/env/generate_random_policies.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"0428bfbecb501b40"}},{"code_sha256_prefix":"f8de7019770bfdbe","entry":"is_in_blocked_area","repo":"Stephanehk/Learning-OA-From-Prefs","repo_kind":"official","path":"learn_advantage/env/generate_random_mdp.py","file_url":"https://github.com/Stephanehk/Learning-OA-From-Prefs/blob/HEAD/learn_advantage/env/generate_random_mdp.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"f8de7019770bfdbe"}},{"code_sha256_prefix":"dcf3a8b47b339ebf","entry":"is_in_gated_area","repo":"Stephanehk/Learning-OA-From-Prefs","repo_kind":"official","path":"learn_advantage/env/generate_random_mdp.py","file_url":"https://github.com/Stephanehk/Learning-OA-From-Prefs/blob/HEAD/learn_advantage/env/generate_random_mdp.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"dcf3a8b47b339ebf"}},{"code_sha256_prefix":"9b36c07728845fbc","entry":"reward_pred_loss","repo":"Stephanehk/Learning-OA-From-Prefs","repo_kind":"official","path":"learn_advantage/algorithms/advantage_learning.py","file_url":"https://github.com/Stephanehk/Learning-OA-From-Prefs/blob/HEAD/learn_advantage/algorithms/advantage_learning.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"9b36c07728845fbc"}},{"code_sha256_prefix":"dd91e5da29ece6a8","entry":"run_single_set","repo":"Stephanehk/Learning-OA-From-Prefs","repo_kind":"official","path":"learn_advantage/algorithms/advantage_learning.py","file_url":"https://github.com/Stephanehk/Learning-OA-From-Prefs/blob/HEAD/learn_advantage/algorithms/advantage_learning.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"dd91e5da29ece6a8"}},{"code_sha256_prefix":"2708f9e87af72605","entry":"parse_args","repo":"Stephanehk/Learning-OA-From-Prefs","repo_kind":"official","path":"learn_advantage/utils/argparse_utils.py","file_url":"https://github.com/Stephanehk/Learning-OA-From-Prefs/blob/HEAD/learn_advantage/utils/argparse_utils.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"2708f9e87af72605"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}