{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/how-to-evaluate-reward-models-for-rlhf","title":"How to Evaluate Reward Models for RLHF","arxiv_id":"2410.14872","date":"2024-10-18","proceeding":null,"authors":["Evan Frick","Tianle Li","Connor Chen","Wei-Lin Chiang","Anastasios N. Angelopoulos","Jiantao Jiao","Banghua Zhu","Joseph E. Gonzalez","Ion Stoica"],"abstract":"We introduce a new benchmark for reward models that quantifies their ability to produce strong language models through RLHF (Reinforcement Learning from Human Feedback). The gold-standard approach is to run a full RLHF training pipeline and directly probe downstream LLM performance. However, this process is prohibitively expensive. To address this, we build a predictive model of downstream LLM performance by evaluating the reward model on proxy tasks. These proxy tasks consist of a large-scale human preference and a verifiable correctness preference dataset, in which we measure 12 metrics across 12 domains. To investigate which reward model metrics are most correlated to gold-standard RLHF outcomes, we launch an end-to-end RLHF experiment on a large-scale crowdsourced human preference platform to view real reward model downstream performance as ground truth. Ultimately, we compile our data and findings into Preference Proxy Evaluations (PPE), the first reward model benchmark explicitly linked to post-RLHF real-world human preference performance, which we open-source for public use and further development. Our code and evaluations can be found at https://github.com/lmarena/PPE .","url_abs":"https://arxiv.org/abs/2410.14872v2","url_pdf":"https://arxiv.org/pdf/2410.14872v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"how-to-evaluate-reward-models-for-rlhf","repo_url":"https://github.com/lmarena/ppe","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=2410.14872","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2410.14872"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/lmarena/ppe","reach":{"status":"ok"}}],"summary":{"ran":6,"ran_draft_wrong":1},"by_repo_kind":{"official":{"samples":7,"ran":7,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":7,"samples":[{"code_sha256_prefix":"79e87e8e79b02b1f","entry":"allow_llm_judge","repo":"lmarena/ppe","repo_kind":"official","path":"utils/scorers.py","file_url":"https://github.com/lmarena/ppe/blob/HEAD/utils/scorers.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"79e87e8e79b02b1f"}},{"code_sha256_prefix":"162803a87a848f49","entry":"chat_completion_openai","repo":"lmarena/ppe","repo_kind":"official","path":"utils/core.py","file_url":"https://github.com/lmarena/ppe/blob/HEAD/utils/core.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"162803a87a848f49"}},{"code_sha256_prefix":"afbaf46000344a52","entry":"contains_list","repo":"lmarena/ppe","repo_kind":"official","path":"display.py","file_url":"https://github.com/lmarena/ppe/blob/HEAD/display.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"afbaf46000344a52"}},{"code_sha256_prefix":"d59fa2d0990791a0","entry":"get_accuracy","repo":"lmarena/ppe","repo_kind":"official","path":"utils/scorers.py","file_url":"https://github.com/lmarena/ppe/blob/HEAD/utils/scorers.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"d59fa2d0990791a0"}},{"code_sha256_prefix":"07e6f41cf857a76f","entry":"make_config","repo":"lmarena/ppe","repo_kind":"official","path":"utils/core.py","file_url":"https://github.com/lmarena/ppe/blob/HEAD/utils/core.py","link_basis":"plan_row","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"07e6f41cf857a76f"}},{"code_sha256_prefix":"4146a53b035f7dbf","entry":"recursive_union","repo":"lmarena/ppe","repo_kind":"official","path":"score.py","file_url":"https://github.com/lmarena/ppe/blob/HEAD/score.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"4146a53b035f7dbf"}},{"code_sha256_prefix":"0289a0e6c43b86d3","entry":"register","repo":"lmarena/ppe","repo_kind":"official","path":"utils/core.py","file_url":"https://github.com/lmarena/ppe/blob/HEAD/utils/core.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"0289a0e6c43b86d3"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}