{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/procedural-dilemma-generation-for-evaluating","title":"Procedural Dilemma Generation for Evaluating Moral Reasoning in Humans and Language Models","arxiv_id":"2404.10975","date":"2024-04-17","proceeding":null,"authors":["Jan-Philipp Fränken","Kanishk Gandhi","Tori Qiu","Ayesha Khawaja","Noah D. Goodman","Tobias Gerstenberg"],"abstract":"As AI systems like language models are increasingly integrated into decision-making processes affecting people's lives, it's critical to ensure that these systems have sound moral reasoning. To test whether they do, we need to develop systematic evaluations. We provide a framework that uses a language model to translate causal graphs that capture key aspects of moral dilemmas into prompt templates. With this framework, we procedurally generated a large and diverse set of moral dilemmas -- the OffTheRails benchmark -- consisting of 50 scenarios and 400 unique test items. We collected moral permissibility and intention judgments from human participants for a subset of our items and compared these judgments to those from two language models (GPT-4 and Claude-2) across eight conditions. We find that moral dilemmas in which the harm is a necessary means (as compared to a side effect) resulted in lower permissibility and higher intention ratings for both participants and language models. The same pattern was observed for evitable versus inevitable harmful outcomes. However, there was no clear effect of whether the harm resulted from an agent's action versus from having omitted to act. We discuss limitations of our prompt generation pipeline and opportunities for improving scenarios to increase the strength of experimental effects.","url_abs":"https://arxiv.org/abs/2404.10975v1","url_pdf":"https://arxiv.org/pdf/2404.10975v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"procedural-dilemma-generation-for-evaluating","repo_url":"https://github.com/cicl-stanford/moral-evals","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"decision-making","task_name":"Decision Making"},{"task_slug":"language-modelling","task_name":"Language Modelling"},{"task_slug":"moral-permissibility","task_name":"Moral Permissibility"}],"methods":[{"method_slug":"set","method_name":"SET"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2404.10975","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2404.10975"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/cicl-stanford/moral-evals","reach":{"status":"ok","spdx":"MIT"}}],"summary":{"ran":5,"unverified":2},"by_repo_kind":{"official":{"samples":7,"ran":5,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"e0fd89f5e55b1a4f","entry":"get_context","repo":"cicl-stanford/moral-evals","repo_kind":"official","path":"offtherails/src/helpers.py","file_url":"https://github.com/cicl-stanford/moral-evals/blob/HEAD/offtherails/src/helpers.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"e0fd89f5e55b1a4f"}},{"code_sha256_prefix":"a58aba4c1171a579","entry":"get_context","repo":"cicl-stanford/moral-evals","repo_kind":"official","path":"offtherails/src/stage_1.py","file_url":"https://github.com/cicl-stanford/moral-evals/blob/HEAD/offtherails/src/stage_1.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"a58aba4c1171a579"}},{"code_sha256_prefix":"4406ad8261c56fb3","entry":"get_example","repo":"cicl-stanford/moral-evals","repo_kind":"official","path":"offtherails/src/helpers.py","file_url":"https://github.com/cicl-stanford/moral-evals/blob/HEAD/offtherails/src/helpers.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"4406ad8261c56fb3"}},{"code_sha256_prefix":"50decf247ea8cfa3","entry":"get_vars_from_out","repo":"cicl-stanford/moral-evals","repo_kind":"official","path":"offtherails/src/helpers.py","file_url":"https://github.com/cicl-stanford/moral-evals/blob/HEAD/offtherails/src/helpers.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"50decf247ea8cfa3"}},{"code_sha256_prefix":"0e5309d0f2137cd8","entry":"parse_response","repo":"cicl-stanford/moral-evals","repo_kind":"official","path":"offtherails/src/evaluate_llm.py","file_url":"https://github.com/cicl-stanford/moral-evals/blob/HEAD/offtherails/src/evaluate_llm.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"0e5309d0f2137cd8"}},{"code_sha256_prefix":"bd4690a435f97ce8","entry":"get_completions","repo":"cicl-stanford/moral-evals","repo_kind":"official","path":"offtherails/src/generate_conditions.py","file_url":"https://github.com/cicl-stanford/moral-evals/blob/HEAD/offtherails/src/generate_conditions.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"bd4690a435f97ce8"}},{"code_sha256_prefix":"96407f5816599b96","entry":"get_example","repo":"cicl-stanford/moral-evals","repo_kind":"official","path":"offtherails/src/stage_1.py","file_url":"https://github.com/cicl-stanford/moral-evals/blob/HEAD/offtherails/src/stage_1.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"96407f5816599b96"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}