{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/unintended-harms-of-value-aligned-llms","title":"Unintended Harms of Value-Aligned LLMs: Psychological and Empirical Insights","arxiv_id":"2506.06404","date":"2025-06-06","proceeding":null,"authors":["Sooyung Choi","JaeHyeok Lee","Xiaoyuan Yi","Jing Yao","Xing Xie","JinYeong Bak"],"abstract":"The application scope of Large Language Models (LLMs) continues to expand, leading to increasing interest in personalized LLMs that align with human values. However, aligning these models with individual values raises significant safety concerns, as certain values may correlate with harmful information. In this paper, we identify specific safety risks associated with value-aligned LLMs and investigate the psychological principles behind these challenges. Our findings reveal two key insights. (1) Value-aligned LLMs are more prone to harmful behavior compared to non-fine-tuned models and exhibit slightly higher risks in traditional safety evaluations than other fine-tuned models. (2) These safety issues arise because value-aligned LLMs genuinely generate text according to the aligned values, which can amplify harmful outcomes. Using a dataset with detailed safety categories, we find significant correlations between value alignment and safety risks, supported by psychological hypotheses. This study offers insights into the \"black box\" of value alignment and proposes in-context alignment methods to enhance the safety of value-aligned LLMs.","url_abs":"https://arxiv.org/abs/2506.06404v1","url_pdf":"https://arxiv.org/pdf/2506.06404v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"unintended-harms-of-value-aligned-llms","repo_url":"https://github.com/Human-Language-Intelligence/Unintended-Harms-LLM","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[],"methods":[{"method_slug":"align","method_name":"ALIGN"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/2506.06404","atlas_url":"https://app.syntology.ai/?focus=2506.06404","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2506.06404"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-25T09:33:49+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/Human-Language-Intelligence/Unintended-Harms-LLM","reach":null},{"provenance":"deterministic:regex_extraction","url":"https://github.com/meta-llama/llama-recipes","reach":{"status":"ok","spdx":"MIT"}}],"summary":{"ran_draft_wrong":1,"unverified":1},"by_repo_kind":{"official":{"samples":1,"ran":1,"repositories":1},"found_in_text":{"samples":1,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":2,"samples":[{"code_sha256_prefix":"0b48977d846ad894","entry":"remove_non_english_characters","repo":"Human-Language-Intelligence/Unintended-Harms-LLM","repo_kind":"official","path":"evaluate/eval_RTP.py","file_url":"https://github.com/Human-Language-Intelligence/Unintended-Harms-LLM/blob/HEAD/evaluate/eval_RTP.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"0b48977d846ad894"}},{"code_sha256_prefix":"e6104bf4c7a58936","entry":"get_llm_response","repo":"meta-llama/llama-recipes","repo_kind":"found_in_text","path":"end-to-end-use-cases/whatsapp_llama_4_bot/ec2_services.py","file_url":"https://github.com/meta-llama/llama-recipes/blob/HEAD/end-to-end-use-cases/whatsapp_llama_4_bot/ec2_services.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":false,"mcp_get_code":{"code_sha256":"e6104bf4c7a58936"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}