{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/robust-multi-objective-controlled-decoding-of","title":"Robust Multi-Objective Controlled Decoding of Large Language Models","arxiv_id":"2503.08796","date":"2025-03-11","proceeding":null,"authors":["Seongho Son","William Bankes","Sangwoong Yoon","Shyam Sundhar Ramesh","Xiaohang Tang","Ilija Bogunovic"],"abstract":"Test-time alignment of Large Language Models (LLMs) to human preferences offers a flexible way to generate responses aligned to diverse objectives without extensive retraining of LLMs. Existing methods achieve alignment to multiple objectives simultaneously (e.g., instruction-following, helpfulness, conciseness) by optimizing their corresponding reward functions. However, they often rely on predefined weights or optimize for averages, sacrificing one objective for another and leading to unbalanced outcomes. To address this, we introduce Robust Multi-Objective Decoding (RMOD), a novel inference-time algorithm that optimizes for improving worst-case rewards. RMOD formalizes the robust decoding problem as a maximin two-player game between reward weights and the sampling policy, solving for the Nash equilibrium. We show that the game reduces to a convex optimization problem to find the worst-case weights, while the best response policy can be computed analytically. We also introduce a practical RMOD variant designed for efficient decoding with contemporary LLMs, incurring minimal computational overhead compared to non-robust Multi-Objective Decoding (MOD) methods. Our experimental results showcase the effectiveness of RMOD in generating responses equitably aligned with diverse objectives, outperforming baselines up to 20%.","url_abs":"https://arxiv.org/abs/2503.08796v1","url_pdf":"https://arxiv.org/pdf/2503.08796v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"robust-multi-objective-controlled-decoding-of","repo_url":"https://github.com/williambankes/robust-multi-objective-decoding","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"instruction-following","task_name":"Instruction Following"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2503.08796","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2503.08796"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/williambankes/robust-multi-objective-decoding","reach":null}],"summary":{"ran":2},"by_repo_kind":{"official":{"samples":2,"ran":2,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":2,"samples":[{"code_sha256_prefix":"2fe809f1612872c6","entry":"BaseMultiObjectiveValueFunction","repo":"williambankes/robust-multi-objective-decoding","repo_kind":"official","path":"robust_multi_objective_decoding/decoders/blockwise_robust_decoder.py","file_url":"https://github.com/williambankes/robust-multi-objective-decoding/blob/HEAD/robust_multi_objective_decoding/decoders/blockwise_robust_decoder.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":"deterministic","behaviour_fingerprint":true,"licence":"GPL-3.0","inline_ok":false,"mcp_get_code":{"code_sha256":"2fe809f1612872c6"}},{"code_sha256_prefix":"36363ceba7dbd2e7","entry":"BlockwiseRobustDecoder","repo":"williambankes/robust-multi-objective-decoding","repo_kind":"official","path":"robust_multi_objective_decoding/decoders/blockwise_robust_decoder.py","file_url":"https://github.com/williambankes/robust-multi-objective-decoding/blob/HEAD/robust_multi_objective_decoding/decoders/blockwise_robust_decoder.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"GPL-3.0","inline_ok":false,"mcp_get_code":{"code_sha256":"36363ceba7dbd2e7"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}