{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/rmb-comprehensively-benchmarking-reward","title":"RMB: Comprehensively Benchmarking Reward Models in LLM Alignment","arxiv_id":"2410.09893","date":"2024-10-13","proceeding":null,"authors":["Enyu Zhou","Guodong Zheng","Binghai Wang","Zhiheng Xi","Shihan Dou","Rong Bao","Wei Shen","Limao Xiong","Jessica Fan","Yurong Mou","Rui Zheng","Tao Gui","Qi Zhang","Xuanjing Huang"],"abstract":"Reward models (RMs) guide the alignment of large language models (LLMs), steering them toward behaviors preferred by humans. Evaluating RMs is the key to better aligning LLMs. However, the current evaluation of RMs may not directly correspond to their alignment performance due to the limited distribution of evaluation data and evaluation methods that are not closely related to alignment objectives. To address these limitations, we propose RMB, a comprehensive RM benchmark that covers over 49 real-world scenarios and includes both pairwise and Best-of-N (BoN) evaluations to better reflect the effectiveness of RMs in guiding alignment optimization. We demonstrate a positive correlation between our benchmark and the downstream alignment task performance. Based on our benchmark, we conduct extensive analysis on the state-of-the-art RMs, revealing their generalization defects that were not discovered by previous benchmarks, and highlighting the potential of generative RMs. Furthermore, we delve into open questions in reward models, specifically examining the effectiveness of majority voting for the evaluation of reward models and analyzing the impact factors of generative RMs, including the influence of evaluation criteria and instructing methods. Our evaluation code and datasets are available at https://github.com/Zhou-Zoey/RMB-Reward-Model-Benchmark.","url_abs":"https://arxiv.org/abs/2410.09893v1","url_pdf":"https://arxiv.org/pdf/2410.09893v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"rmb-comprehensively-benchmarking-reward","repo_url":"https://github.com/zhou-zoey/rmb-reward-model-benchmark","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"benchmarking","task_name":"Benchmarking"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/2410.09893","atlas_url":"https://app.syntology.ai/?focus=2410.09893","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2410.09893"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-25T09:33:49+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/zhou-zoey/rmb-reward-model-benchmark","reach":null}],"summary":{"ran_draft_wrong":2,"ran_honours":1},"by_repo_kind":{"official":{"samples":2,"ran":2,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":3,"samples":[{"code_sha256_prefix":"4cb00c8ef6768ef4","entry":"load_config","repo":null,"repo_kind":null,"path":null,"file_url":null,"link_basis":"identical_code_first_harvested_elsewhere","language":"python","status":"ran_draft_wrong","verification_level":2,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":null,"inline_ok":false,"mcp_get_code":{"code_sha256":"4cb00c8ef6768ef4"}},{"code_sha256_prefix":"e95b318087b7ce83","entry":"find_json_files","repo":"zhou-zoey/rmb-reward-model-benchmark","repo_kind":"official","path":"eval/scripts/my_run_rm.py","file_url":"https://github.com/zhou-zoey/rmb-reward-model-benchmark/blob/HEAD/eval/scripts/my_run_rm.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"e95b318087b7ce83"}},{"code_sha256_prefix":"ad14d89d1d8e6f44","entry":"get_parameters","repo":"zhou-zoey/rmb-reward-model-benchmark","repo_kind":"official","path":"eval/scripts/my_run_rm.py","file_url":"https://github.com/zhou-zoey/rmb-reward-model-benchmark/blob/HEAD/eval/scripts/my_run_rm.py","link_basis":"first_harvest_node","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":"well_formed","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"ad14d89d1d8e6f44"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}