{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/fine-tuning-language-models-with-advantage","title":"Fine-Tuning Language Models with Advantage-Induced Policy Alignment","arxiv_id":"2306.02231","date":"2023-06-04","proceeding":null,"authors":["Banghua Zhu","Hiteshi Sharma","Felipe Vieira Frujeri","Shi Dong","Chenguang Zhu","Michael I. Jordan","Jiantao Jiao"],"abstract":"Reinforcement learning from human feedback (RLHF) has emerged as a reliable approach to aligning large language models (LLMs) to human preferences. Among the plethora of RLHF techniques, proximal policy optimization (PPO) is of the most widely used methods. Despite its popularity, however, PPO may suffer from mode collapse, instability, and poor sample efficiency. We show that these issues can be alleviated by a novel algorithm that we refer to as Advantage-Induced Policy Alignment (APA), which leverages a squared error loss function based on the estimated advantages. We demonstrate empirically that APA consistently outperforms PPO in language tasks by a large margin, when a separate reward model is employed as the evaluator. In addition, compared with PPO, APA offers a more stable form of control over the deviation from the model's initial policy, ensuring that the model improves its performance without collapsing to deterministic output. In addition to empirical results, we also provide a theoretical justification supporting the design of our loss function.","url_abs":"https://arxiv.org/abs/2306.02231v3","url_pdf":"https://arxiv.org/pdf/2306.02231v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"fine-tuning-language-models-with-advantage","repo_url":"https://github.com/microsoft/rlhf-apa","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[],"methods":[{"method_slug":"apa","method_name":"APA"},{"method_slug":"entropy-regularization","method_name":"Entropy Regularization"},{"method_slug":"ppo","method_name":"PPO"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2306.02231","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2306.02231"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/microsoft/rlhf-apa","reach":{"status":"ok","spdx":"MIT"}}],"summary":{"ran":4,"unverified":3},"by_repo_kind":{"official":{"samples":7,"ran":4,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"2ab02dee00149590","entry":"make_head","repo":"microsoft/rlhf-apa","repo_kind":"official","path":"trlx/utils/modeling.py","file_url":"https://github.com/microsoft/rlhf-apa/blob/HEAD/trlx/utils/modeling.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"2ab02dee00149590"}},{"code_sha256_prefix":"e2c3cae9e92c72e8","entry":"parse_result","repo":"microsoft/rlhf-apa","repo_kind":"official","path":"trlx/ray_tune/wandb.py","file_url":"https://github.com/microsoft/rlhf-apa/blob/HEAD/trlx/ray_tune/wandb.py","link_basis":"plan_row","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"e2c3cae9e92c72e8"}},{"code_sha256_prefix":"ace162af010af5d8","entry":"rhasattr","repo":"microsoft/rlhf-apa","repo_kind":"official","path":"trlx/utils/modeling.py","file_url":"https://github.com/microsoft/rlhf-apa/blob/HEAD/trlx/utils/modeling.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"ace162af010af5d8"}},{"code_sha256_prefix":"3bfb7deb629cb756","entry":"topk_mask","repo":"microsoft/rlhf-apa","repo_kind":"official","path":"trlx/models/modeling_ilql.py","file_url":"https://github.com/microsoft/rlhf-apa/blob/HEAD/trlx/models/modeling_ilql.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"3bfb7deb629cb756"}},{"code_sha256_prefix":"c80bffc6c3b486fc","entry":"create_report","repo":"microsoft/rlhf-apa","repo_kind":"official","path":"trlx/ray_tune/wandb.py","file_url":"https://github.com/microsoft/rlhf-apa/blob/HEAD/trlx/ray_tune/wandb.py","link_basis":"plan_row","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"c80bffc6c3b486fc"}},{"code_sha256_prefix":"fc2499cfa7162602","entry":"rgetattr","repo":"microsoft/rlhf-apa","repo_kind":"official","path":"trlx/utils/modeling.py","file_url":"https://github.com/microsoft/rlhf-apa/blob/HEAD/trlx/utils/modeling.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"fc2499cfa7162602"}},{"code_sha256_prefix":"6729b7c27552d487","entry":"tokenize_dialogue","repo":"microsoft/rlhf-apa","repo_kind":"official","path":"trlx/pipeline/offline_pipeline.py","file_url":"https://github.com/microsoft/rlhf-apa/blob/HEAD/trlx/pipeline/offline_pipeline.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"6729b7c27552d487"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}