{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/are-large-language-models-consistent-over","title":"Are Large Language Models Consistent over Value-laden Questions?","arxiv_id":"2407.02996","date":"2024-07-03","proceeding":null,"authors":["Jared Moore","Tanvi Deshpande","Diyi Yang"],"abstract":"Large language models (LLMs) appear to bias their survey answers toward certain values. Nonetheless, some argue that LLMs are too inconsistent to simulate particular values. Are they? To answer, we first define value consistency as the similarity of answers across (1) paraphrases of one question, (2) related questions under one topic, (3) multiple-choice and open-ended use-cases of one question, and (4) multilingual translations of a question to English, Chinese, German, and Japanese. We apply these measures to small and large, open LLMs including llama-3, as well as gpt-4o, using 8,000 questions spanning more than 300 topics. Unlike prior work, we find that models are relatively consistent across paraphrases, use-cases, translations, and within a topic. Still, some inconsistencies remain. Models are more consistent on uncontroversial topics (e.g., in the U.S., \"Thanksgiving\") than on controversial ones (\"euthanasia\"). Base models are both more consistent compared to fine-tuned models and are uniform in their consistency across topics, while fine-tuned models are more inconsistent about some topics (\"euthanasia\") than others (\"women's rights\") like our human subjects (n=165).","url_abs":"https://arxiv.org/abs/2407.02996v2","url_pdf":"https://arxiv.org/pdf/2407.02996v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"are-large-language-models-consistent-over","repo_url":"https://github.com/jlcmoore/ValueConsistency","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"multiple-choice","task_name":"Multiple-choice"}],"methods":[{"method_slug":"base","method_name":"BASE"}],"datasets_introduced":[{"slug":"valueconsistency","name":"ValueConsistency","full_name":""}],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2407.02996","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2407.02996"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/jlcmoore/ValueConsistency","reach":null}],"summary":{"ran_draft_wrong":3},"by_repo_kind":{"official":{"samples":3,"ran":3,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"fd0a0608c9784181","entry":"base64_decode","repo":"jlcmoore/ValueConsistency","repo_kind":"official","path":"scripts/format_mturk.py","file_url":"https://github.com/jlcmoore/ValueConsistency/blob/HEAD/scripts/format_mturk.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"fd0a0608c9784181"}},{"code_sha256_prefix":"98152d39ac580e70","entry":"base64_encode","repo":"jlcmoore/ValueConsistency","repo_kind":"official","path":"scripts/format_mturk.py","file_url":"https://github.com/jlcmoore/ValueConsistency/blob/HEAD/scripts/format_mturk.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"98152d39ac580e70"}},{"code_sha256_prefix":"131e438287bb3aa3","entry":"group_n_rows","repo":"jlcmoore/ValueConsistency","repo_kind":"official","path":"scripts/format_mturk.py","file_url":"https://github.com/jlcmoore/ValueConsistency/blob/HEAD/scripts/format_mturk.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"131e438287bb3aa3"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}