{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/what-makes-a-reward-model-a-good-teacher-an","title":"What Makes a Reward Model a Good Teacher? An Optimization Perspective","arxiv_id":"2503.15477","date":"2025-03-19","proceeding":null,"authors":["Noam Razin","Zixuan Wang","Hubert Strauss","Stanley Wei","Jason D. Lee","Sanjeev Arora"],"abstract":"The success of Reinforcement Learning from Human Feedback (RLHF) critically depends on the quality of the reward model. While this quality is primarily evaluated through accuracy, it remains unclear whether accuracy fully captures what makes a reward model an effective teacher. We address this question from an optimization perspective. First, we prove that regardless of how accurate a reward model is, if it induces low reward variance, then the RLHF objective suffers from a flat landscape. Consequently, even a perfectly accurate reward model can lead to extremely slow optimization, underperforming less accurate models that induce higher reward variance. We additionally show that a reward model that works well for one language model can induce low reward variance, and thus a flat objective landscape, for another. These results establish a fundamental limitation of evaluating reward models solely based on accuracy or independently of the language model they guide. Experiments using models of up to 8B parameters corroborate our theory, demonstrating the interplay between reward variance, accuracy, and reward maximization rate. Overall, our findings highlight that beyond accuracy, a reward model needs to induce sufficient variance for efficient optimization.","url_abs":"https://arxiv.org/abs/2503.15477v1","url_pdf":"https://arxiv.org/pdf/2503.15477v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"what-makes-a-reward-model-a-good-teacher-an","repo_url":"https://github.com/princeton-pli/what-makes-good-rm","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"language-modeling","task_name":"Language Modeling"},{"task_slug":"language-modelling","task_name":"Language Modelling"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/2503.15477","atlas_url":"https://app.syntology.ai/?focus=2503.15477","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2503.15477"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-25T09:33:49+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/princeton-pli/what-makes-good-rm","reach":{"status":"ok","spdx":"MIT"}}],"summary":{"ran":4,"unverified":1},"by_repo_kind":{"official":{"samples":5,"ran":4,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"e5e70592012299c3","entry":"convert_config_dict_to_str_args","repo":"princeton-pli/what-makes-good-rm","repo_kind":"official","path":"src/what_makes_good_rm/Arguments/arg_utils.py","file_url":"https://github.com/princeton-pli/what-makes-good-rm/blob/HEAD/src/what_makes_good_rm/Arguments/arg_utils.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"e5e70592012299c3"}},{"code_sha256_prefix":"eb00910c6cc427e5","entry":"create_logger","repo":"princeton-pli/what-makes-good-rm","repo_kind":"official","path":"src/what_makes_good_rm/Utils/single_process_logging.py","file_url":"https://github.com/princeton-pli/what-makes-good-rm/blob/HEAD/src/what_makes_good_rm/Utils/single_process_logging.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"eb00910c6cc427e5"}},{"code_sha256_prefix":"de2925f13f0eb807","entry":"parse_unknown_argparse_args_into_dict","repo":"princeton-pli/what-makes-good-rm","repo_kind":"official","path":"src/what_makes_good_rm/Arguments/arg_utils.py","file_url":"https://github.com/princeton-pli/what-makes-good-rm/blob/HEAD/src/what_makes_good_rm/Arguments/arg_utils.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"de2925f13f0eb807"}},{"code_sha256_prefix":"79deb9297d937d86","entry":"update_model_num_embeddings_and_special_tokens","repo":"princeton-pli/what-makes-good-rm","repo_kind":"official","path":"src/what_makes_good_rm/Utils/sharedmisc.py","file_url":"https://github.com/princeton-pli/what-makes-good-rm/blob/HEAD/src/what_makes_good_rm/Utils/sharedmisc.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"79deb9297d937d86"}},{"code_sha256_prefix":"2188119dc834b8a0","entry":"get_logger","repo":"princeton-pli/what-makes-good-rm","repo_kind":"official","path":"src/what_makes_good_rm/Utils/logger.py","file_url":"https://github.com/princeton-pli/what-makes-good-rm/blob/HEAD/src/what_makes_good_rm/Utils/logger.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"2188119dc834b8a0"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}