{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/bi-factorial-preference-optimization","title":"Bi-Factorial Preference Optimization: Balancing Safety-Helpfulness in Language Models","arxiv_id":"2408.15313","date":"2024-08-27","proceeding":null,"authors":["Wenxuan Zhang","Philip H. S. Torr","Mohamed Elhoseiny","Adel Bibi"],"abstract":"Fine-tuning large language models (LLMs) on human preferences, typically through reinforcement learning from human feedback (RLHF), has proven successful in enhancing their capabilities. However, ensuring the safety of LLMs during the fine-tuning remains a critical concern, and mitigating the potential conflicts in safety and helpfulness is costly in RLHF. To address this issue, we propose a supervised learning framework called Bi-Factorial Preference Optimization (BFPO), which re-parameterizes a joint RLHF objective of both safety and helpfulness into a single supervised learning objective. In the supervised optimization, a labeling function is used to capture global preferences ranking to balance both safety and helpfulness. To evaluate BFPO, we develop a benchmark including comprehensive discriminative and generative tasks for helpfulness and harmlessness. The results indicate that our method significantly outperforms existing approaches in both safety and helpfulness. Moreover, BFPO eliminates the need for human prompting and annotation in LLM fine-tuning while achieving the same level of safety as methods that heavily rely on human labor, with less than 10% of the computational resources. The training recipes and models will be released.","url_abs":"https://arxiv.org/abs/2408.15313v1","url_pdf":"https://arxiv.org/pdf/2408.15313v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2408.15313","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2408.15313"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"deterministic:regex_extraction","url":"https://github.com/wx-zhang/bfpo","reach":{"status":"ok"}}],"summary":{"ran":2,"unverified":1},"by_repo_kind":{"found_in_text":{"samples":3,"ran":2,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":3,"samples":[{"code_sha256_prefix":"4c7a20efd8b515fe","entry":"add_dataset_name","repo":"wx-zhang/bfpo","repo_kind":"found_in_text","path":"src/alignment/data.py","file_url":"https://github.com/wx-zhang/bfpo/blob/HEAD/src/alignment/data.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"4c7a20efd8b515fe"}},{"code_sha256_prefix":"ecf650a71ce5efc6","entry":"apply_chat_template","repo":"wx-zhang/bfpo","repo_kind":"found_in_text","path":"src/alignment/data.py","file_url":"https://github.com/wx-zhang/bfpo/blob/HEAD/src/alignment/data.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"ecf650a71ce5efc6"}},{"code_sha256_prefix":"c9ed692d311c459a","entry":"get_quantization_config","repo":"wx-zhang/bfpo","repo_kind":"found_in_text","path":"src/alignment/model_utils.py","file_url":"https://github.com/wx-zhang/bfpo/blob/HEAD/src/alignment/model_utils.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"c9ed692d311c459a"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}