{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/self-rewarding-language-models","title":"Self-Rewarding Language Models","arxiv_id":"2401.10020","date":"2024-01-18","proceeding":null,"authors":["Weizhe Yuan","Richard Yuanzhe Pang","Kyunghyun Cho","Xian Li","Sainbayar Sukhbaatar","Jing Xu","Jason Weston"],"abstract":"We posit that to achieve superhuman agents, future models require superhuman feedback in order to provide an adequate training signal. Current approaches commonly train reward models from human preferences, which may then be bottlenecked by human performance level, and secondly these separate frozen reward models cannot then learn to improve during LLM training. In this work, we study Self-Rewarding Language Models, where the language model itself is used via LLM-as-a-Judge prompting to provide its own rewards during training. We show that during Iterative DPO training that not only does instruction following ability improve, but also the ability to provide high-quality rewards to itself. Fine-tuning Llama 2 70B on three iterations of our approach yields a model that outperforms many existing systems on the AlpacaEval 2.0 leaderboard, including Claude 2, Gemini Pro, and GPT-4 0613. While there is much left still to explore, this work opens the door to the possibility of models that can continually improve in both axes.","url_abs":"https://arxiv.org/abs/2401.10020v3","url_pdf":"https://arxiv.org/pdf/2401.10020v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"self-rewarding-language-models","repo_url":"https://github.com/lucidrains/self-rewarding-lm-pytorch","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"self-rewarding-language-models","repo_url":"https://github.com/safouaneelg/SRT2I","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"NOASSERTION"}},{"paper_slug":"self-rewarding-language-models","repo_url":"https://github.com/gagan3012/self_rewarding_models","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"pytorch","reach":{"status":"ok","spdx":"Apache-2.0"}}],"tasks":[{"task_slug":"instruction-following","task_name":"Instruction Following"},{"task_slug":"language-modeling","task_name":"Language Modeling"},{"task_slug":"language-modelling","task_name":"Language Modelling"}],"methods":[{"method_slug":"absolute-position-encodings","method_name":"Absolute Position Encodings"},{"method_slug":"adam","method_name":"Adam"},{"method_slug":"attention","method_name":"Attention"},{"method_slug":"bpe","method_name":"BPE"},{"method_slug":"dpo","method_name":"DPO"},{"method_slug":"dense-connections","method_name":"Dense Connections"},{"method_slug":"dropout","method_name":"Dropout"},{"method_slug":"gpt-4","method_name":"GPT-4"},{"method_slug":"label-smoothing","method_name":"Label Smoothing"},{"method_slug":"layer-normalization","method_name":"Layer Normalization"},{"method_slug":"linear-layer","method_name":"Linear Layer"},{"method_slug":"multi-head-attention","method_name":"Multi-Head Attention"},{"method_slug":"position-wise-feed-forward-layer","method_name":"Position-Wise Feed-Forward Layer"},{"method_slug":"residual-connection","method_name":"Residual Connection"},{"method_slug":"softmax","method_name":"Softmax"},{"method_slug":"transformer","method_name":"Transformer"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2401.10020","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2401.10020"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/lucidrains/self-rewarding-lm-pytorch","reach":{"status":"ok","spdx":"MIT"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/gagan3012/self_rewarding_models","reach":{"status":"ok","spdx":"Apache-2.0"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/safouaneelg/SRT2I","reach":{"status":"ok","spdx":"NOASSERTION"}}],"summary":{"ran_draft_wrong":1,"ran_violates":2,"unverified":8},"by_repo_kind":{"listed":{"samples":11,"ran":3,"repositories":2}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"b5472673332de78e","entry":"log","repo":"lucidrains/self-rewarding-lm-pytorch","repo_kind":"listed","path":"self_rewarding_lm_pytorch/sampling_utils.py","file_url":"https://github.com/lucidrains/self-rewarding-lm-pytorch/blob/HEAD/self_rewarding_lm_pytorch/sampling_utils.py","link_basis":"harvester_set","language":"python","status":"ran_draft_wrong","verification_level":2,"contract_check":"MISDECLARED","metamorphic_tier":"invariant","behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"b5472673332de78e"}},{"code_sha256_prefix":"99b3563ebc6734ad","entry":"default","repo":"lucidrains/self-rewarding-lm-pytorch","repo_kind":"listed","path":"self_rewarding_lm_pytorch/dpo.py","file_url":"https://github.com/lucidrains/self-rewarding-lm-pytorch/blob/HEAD/self_rewarding_lm_pytorch/dpo.py","link_basis":"harvester_set","language":"python","status":"ran_violates","verification_level":1,"contract_check":"VIOLATES","metamorphic_tier":"well_formed","behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"99b3563ebc6734ad"}},{"code_sha256_prefix":"b7c3487f192e31b7","entry":"exists","repo":"lucidrains/self-rewarding-lm-pytorch","repo_kind":"listed","path":"self_rewarding_lm_pytorch/dpo.py","file_url":"https://github.com/lucidrains/self-rewarding-lm-pytorch/blob/HEAD/self_rewarding_lm_pytorch/dpo.py","link_basis":"harvester_set","language":"python","status":"ran_violates","verification_level":1,"contract_check":"VIOLATES","metamorphic_tier":"deterministic","behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"b7c3487f192e31b7"}},{"code_sha256_prefix":"55edf480b3f62cac","entry":"alternate_messages","repo":"gagan3012/self_rewarding_models","repo_kind":"listed","path":"data_gen.py","file_url":"https://github.com/gagan3012/self_rewarding_models/blob/HEAD/data_gen.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"55edf480b3f62cac"}},{"code_sha256_prefix":"291c6f7cd3e538f1","entry":"always","repo":"lucidrains/self-rewarding-lm-pytorch","repo_kind":"listed","path":"self_rewarding_lm_pytorch/mocks.py","file_url":"https://github.com/lucidrains/self-rewarding-lm-pytorch/blob/HEAD/self_rewarding_lm_pytorch/mocks.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"291c6f7cd3e538f1"}},{"code_sha256_prefix":"3429aba36f0d19cc","entry":"create_mock_dataset","repo":"lucidrains/self-rewarding-lm-pytorch","repo_kind":"listed","path":"self_rewarding_lm_pytorch/mocks.py","file_url":"https://github.com/lucidrains/self-rewarding-lm-pytorch/blob/HEAD/self_rewarding_lm_pytorch/mocks.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"3429aba36f0d19cc"}},{"code_sha256_prefix":"7eeab84522869fee","entry":"find_variables_from_jinja_template","repo":"lucidrains/self-rewarding-lm-pytorch","repo_kind":"listed","path":"self_rewarding_lm_pytorch/self_rewarding_lm_pytorch.py","file_url":"https://github.com/lucidrains/self-rewarding-lm-pytorch/blob/HEAD/self_rewarding_lm_pytorch/self_rewarding_lm_pytorch.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"7eeab84522869fee"}},{"code_sha256_prefix":"fd771b0e28e33ab9","entry":"fix_prompt","repo":"gagan3012/self_rewarding_models","repo_kind":"listed","path":"reward_train.py","file_url":"https://github.com/gagan3012/self_rewarding_models/blob/HEAD/reward_train.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"fd771b0e28e33ab9"}},{"code_sha256_prefix":"0e980d2d900070e6","entry":"prompt_format","repo":"gagan3012/self_rewarding_models","repo_kind":"listed","path":"rank_responses.py","file_url":"https://github.com/gagan3012/self_rewarding_models/blob/HEAD/rank_responses.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"0e980d2d900070e6"}},{"code_sha256_prefix":"60cc0acbbea98eb4","entry":"prompt_mask_from_len","repo":"lucidrains/self-rewarding-lm-pytorch","repo_kind":"listed","path":"self_rewarding_lm_pytorch/spin.py","file_url":"https://github.com/lucidrains/self-rewarding-lm-pytorch/blob/HEAD/self_rewarding_lm_pytorch/spin.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"60cc0acbbea98eb4"}},{"code_sha256_prefix":"8b9058a64b018063","entry":"setup_eft","repo":"gagan3012/self_rewarding_models","repo_kind":"listed","path":"data_gen.py","file_url":"https://github.com/gagan3012/self_rewarding_models/blob/HEAD/data_gen.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"8b9058a64b018063"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}