{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/from-demonstrations-to-rewards-alignment","title":"From Demonstrations to Rewards: Alignment Without Explicit Human Preferences","arxiv_id":"2503.13538","date":"2025-03-15","proceeding":null,"authors":["Siliang Zeng","Yao Liu","Huzefa Rangwala","George Karypis","Mingyi Hong","Rasool Fakoor"],"abstract":"One of the challenges of aligning large models with human preferences lies in both the data requirements and the technical complexities of current approaches. Predominant methods, such as RLHF, involve multiple steps, each demanding distinct types of data, including demonstration data and preference data. In RLHF, human preferences are typically modeled through a reward model, which serves as a proxy to guide policy learning during the reinforcement learning stage, ultimately producing a policy aligned with human preferences. However, in this paper, we propose a fresh perspective on learning alignment based on inverse reinforcement learning principles, where the optimal policy is still derived from reward maximization. However, instead of relying on preference data, we directly learn the reward model from demonstration data. This new formulation offers the flexibility to be applied even when only demonstration data is available, a capability that current RLHF methods lack, and it also shows that demonstration data offers more utility than what conventional wisdom suggests. Our extensive evaluation, based on public reward benchmark, HuggingFace Open LLM Leaderboard and MT-Bench, demonstrates that our approach compares favorably to state-of-the-art methods that rely solely on demonstration data.","url_abs":"https://arxiv.org/abs/2503.13538v1","url_pdf":"https://arxiv.org/pdf/2503.13538v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"from-demonstrations-to-rewards-alignment","repo_url":"https://github.com/Hong-Lab-UMN-ECE/IRLAlignment","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=2503.13538","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2503.13538"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/Hong-Lab-UMN-ECE/IRLAlignment","reach":{"status":"ok","spdx":"MIT"}}],"summary":{"ran":7,"ran_draft_wrong":1,"unverified":2},"by_repo_kind":{"official":{"samples":10,"ran":8,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"f99955c55f25123e","entry":"ceil_div","repo":"Hong-Lab-UMN-ECE/IRLAlignment","repo_kind":"official","path":"summarize_from_feedback_details/tldr_dataset.py","file_url":"https://github.com/Hong-Lab-UMN-ECE/IRLAlignment/blob/HEAD/summarize_from_feedback_details/tldr_dataset.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"f99955c55f25123e"}},{"code_sha256_prefix":"579f0202bb30e003","entry":"first_true_indices","repo":"Hong-Lab-UMN-ECE/IRLAlignment","repo_kind":"official","path":"summarize_from_feedback_details/IRL_reward.py","file_url":"https://github.com/Hong-Lab-UMN-ECE/IRLAlignment/blob/HEAD/summarize_from_feedback_details/IRL_reward.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"579f0202bb30e003"}},{"code_sha256_prefix":"1f9a50a3becd157e","entry":"forward","repo":"Hong-Lab-UMN-ECE/IRLAlignment","repo_kind":"official","path":"summarize_from_feedback_details/dpo.py","file_url":"https://github.com/Hong-Lab-UMN-ECE/IRLAlignment/blob/HEAD/summarize_from_feedback_details/dpo.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"1f9a50a3becd157e"}},{"code_sha256_prefix":"51062aca9db3f84c","entry":"generate","repo":"Hong-Lab-UMN-ECE/IRLAlignment","repo_kind":"official","path":"summarize_from_feedback_details/dpo.py","file_url":"https://github.com/Hong-Lab-UMN-ECE/IRLAlignment/blob/HEAD/summarize_from_feedback_details/dpo.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"51062aca9db3f84c"}},{"code_sha256_prefix":"9ad5922df265477f","entry":"layer_init","repo":"Hong-Lab-UMN-ECE/IRLAlignment","repo_kind":"official","path":"visualize_tokens.py","file_url":"https://github.com/Hong-Lab-UMN-ECE/IRLAlignment/blob/HEAD/visualize_tokens.py","link_basis":"harvester_set","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"9ad5922df265477f"}},{"code_sha256_prefix":"16655f3447355899","entry":"masked_mean","repo":"Hong-Lab-UMN-ECE/IRLAlignment","repo_kind":"official","path":"summarize_from_feedback_details/data_pairing.py","file_url":"https://github.com/Hong-Lab-UMN-ECE/IRLAlignment/blob/HEAD/summarize_from_feedback_details/data_pairing.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"16655f3447355899"}},{"code_sha256_prefix":"3a136e5f5a8047ac","entry":"masked_var","repo":"Hong-Lab-UMN-ECE/IRLAlignment","repo_kind":"official","path":"summarize_from_feedback_details/data_pairing.py","file_url":"https://github.com/Hong-Lab-UMN-ECE/IRLAlignment/blob/HEAD/summarize_from_feedback_details/data_pairing.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"3a136e5f5a8047ac"}},{"code_sha256_prefix":"e2429ac73eafe86e","entry":"truncate_response","repo":"Hong-Lab-UMN-ECE/IRLAlignment","repo_kind":"official","path":"summarize_from_feedback_details/sft.py","file_url":"https://github.com/Hong-Lab-UMN-ECE/IRLAlignment/blob/HEAD/summarize_from_feedback_details/sft.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"e2429ac73eafe86e"}},{"code_sha256_prefix":"c64a8d32a7ed56ca","entry":"get_reward","repo":"Hong-Lab-UMN-ECE/IRLAlignment","repo_kind":"official","path":"summarize_from_feedback_details/IRL_reward.py","file_url":"https://github.com/Hong-Lab-UMN-ECE/IRLAlignment/blob/HEAD/summarize_from_feedback_details/IRL_reward.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"c64a8d32a7ed56ca"}},{"code_sha256_prefix":"3dcde3c8b6d6c204","entry":"process_query","repo":"Hong-Lab-UMN-ECE/IRLAlignment","repo_kind":"official","path":"summarize_from_feedback_details/tldr_dataset.py","file_url":"https://github.com/Hong-Lab-UMN-ECE/IRLAlignment/blob/HEAD/summarize_from_feedback_details/tldr_dataset.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"3dcde3c8b6d6c204"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}