{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/mpo-multilingual-safety-alignment-via-reward","title":"MPO: Multilingual Safety Alignment via Reward Gap Optimization","arxiv_id":"2505.16869","date":"2025-05-22","proceeding":null,"authors":["Weixiang Zhao","Yulin Hu","Yang Deng","Tongtong Wu","Wenxuan Zhang","Jiahe Guo","An Zhang","Yanyan Zhao","Bing Qin","Tat-Seng Chua","Ting Liu"],"abstract":"Large language models (LLMs) have become increasingly central to AI applications worldwide, necessitating robust multilingual safety alignment to ensure secure deployment across diverse linguistic contexts. Existing preference learning methods for safety alignment, such as RLHF and DPO, are primarily monolingual and struggle with noisy multilingual data. To address these limitations, we introduce Multilingual reward gaP Optimization (MPO), a novel approach that leverages the well-aligned safety capabilities of the dominant language (English) to improve safety alignment across multiple languages. MPO directly minimizes the reward gap difference between the dominant language and target languages, effectively transferring safety capabilities while preserving the original strengths of the dominant language. Extensive experiments on three LLMs, LLaMA-3.1, Gemma-2 and Qwen2.5, validate MPO's efficacy in multilingual safety alignment without degrading general multilingual utility.","url_abs":"https://arxiv.org/abs/2505.16869v1","url_pdf":"https://arxiv.org/pdf/2505.16869v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"mpo-multilingual-safety-alignment-via-reward","repo_url":"https://github.com/circle-hit/mpo","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":{"status":"ok","spdx":"Apache-2.0"}}],"tasks":[{"task_slug":"safety-alignment","task_name":"Safety Alignment"}],"methods":[{"method_slug":"dpo","method_name":"DPO"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/2505.16869","atlas_url":"https://app.syntology.ai/?focus=2505.16869","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2505.16869"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-25T09:33:49+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"deterministic:regex_extraction","url":"https://github.com/circle-hit/MPO","reach":{"status":"ok","spdx":"Apache-2.0"}}],"summary":{"ran":2,"unverified":1},"by_repo_kind":{"official":{"samples":3,"ran":2,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"f31c69421a40b0a4","entry":"create_score_evaluation_response","repo":"circle-hit/MPO","repo_kind":"official","path":"src/llamafactory/api/chat.py","file_url":"https://github.com/circle-hit/MPO/blob/HEAD/src/llamafactory/api/chat.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"f31c69421a40b0a4"}},{"code_sha256_prefix":"57583cf880d95444","entry":"jsonify","repo":"circle-hit/MPO","repo_kind":"official","path":"src/llamafactory/api/common.py","file_url":"https://github.com/circle-hit/MPO/blob/HEAD/src/llamafactory/api/common.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"57583cf880d95444"}},{"code_sha256_prefix":"87aa9acaa81294f5","entry":"dictify","repo":"circle-hit/MPO","repo_kind":"official","path":"src/llamafactory/api/common.py","file_url":"https://github.com/circle-hit/MPO/blob/HEAD/src/llamafactory/api/common.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"87aa9acaa81294f5"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}