{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/evaluating-the-moral-beliefs-encoded-in-llms-1","title":"Evaluating the Moral Beliefs Encoded in LLMs","arxiv_id":"2307.14324","date":"2023-07-26","proceeding":"NeurIPS 2023 11","authors":["Nino Scherrer","Claudia Shi","Amir Feder","David M. Blei"],"abstract":"This paper presents a case study on the design, administration, post-processing, and evaluation of surveys on large language models (LLMs). It comprises two components: (1) A statistical method for eliciting beliefs encoded in LLMs. We introduce statistical measures and evaluation metrics that quantify the probability of an LLM \"making a choice\", the associated uncertainty, and the consistency of that choice. (2) We apply this method to study what moral beliefs are encoded in different LLMs, especially in ambiguous cases where the right choice is not obvious. We design a large-scale survey comprising 680 high-ambiguity moral scenarios (e.g., \"Should I tell a white lie?\") and 687 low-ambiguity moral scenarios (e.g., \"Should I stop for a pedestrian on the road?\"). Each scenario includes a description, two possible actions, and auxiliary labels indicating violated rules (e.g., \"do not kill\"). We administer the survey to 28 open- and closed-source LLMs. We find that (a) in unambiguous scenarios, most models \"choose\" actions that align with commonsense. In ambiguous cases, most models express uncertainty. (b) Some models are uncertain about choosing the commonsense action because their responses are sensitive to the question-wording. (c) Some models reflect clear preferences in ambiguous scenarios. Specifically, closed-source models tend to agree with each other.","url_abs":"https://arxiv.org/abs/2307.14324v1","url_pdf":"https://arxiv.org/pdf/2307.14324v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"evaluating-the-moral-beliefs-encoded-in-llms-1","repo_url":"https://github.com/ninodimontalcino/moralchoice","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"moral-scenarios","task_name":"Moral Scenarios"},{"task_slug":"survey","task_name":"Survey"}],"methods":[{"method_slug":"align","method_name":"ALIGN"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/2307.14324","atlas_url":"https://app.syntology.ai/?focus=2307.14324","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2307.14324"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-25T09:33:49+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/ninodimontalcino/moralchoice","reach":{"status":"ok","spdx":"MIT"}}],"summary":{"ran":1,"unverified":3},"by_repo_kind":{"official":{"samples":4,"ran":1,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"b2f45a46db40dd60","entry":"get_dict_from_string","repo":"ninodimontalcino/moralchoice","repo_kind":"official","path":"src/utils.py","file_url":"https://github.com/ninodimontalcino/moralchoice/blob/HEAD/src/utils.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"b2f45a46db40dd60"}},{"code_sha256_prefix":"6b031bfdf5694681","entry":"create_model","repo":"ninodimontalcino/moralchoice","repo_kind":"official","path":"src/models.py","file_url":"https://github.com/ninodimontalcino/moralchoice/blob/HEAD/src/models.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"6b031bfdf5694681"}},{"code_sha256_prefix":"9454c616d3d7d219","entry":"get_raw_likelihoods_from_answer","repo":"ninodimontalcino/moralchoice","repo_kind":"official","path":"src/models.py","file_url":"https://github.com/ninodimontalcino/moralchoice/blob/HEAD/src/models.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"9454c616d3d7d219"}},{"code_sha256_prefix":"86958804201c8703","entry":"stem_sentences","repo":"ninodimontalcino/moralchoice","repo_kind":"official","path":"src/utils.py","file_url":"https://github.com/ninodimontalcino/moralchoice/blob/HEAD/src/utils.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"86958804201c8703"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}