{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/do-language-models-exhibit-the-same-cognitive","title":"Do Language Models Exhibit the Same Cognitive Biases in Problem Solving as Human Learners?","arxiv_id":"2401.18070","date":"2024-01-31","proceeding":null,"authors":["Andreas Opedal","Alessandro Stolfo","Haruki Shirakami","Ying Jiao","Ryan Cotterell","Bernhard Schölkopf","Abulhair Saparov","Mrinmaya Sachan"],"abstract":"There is increasing interest in employing large language models (LLMs) as cognitive models. For such purposes, it is central to understand which properties of human cognition are well-modeled by LLMs, and which are not. In this work, we study the biases of LLMs in relation to those known in children when solving arithmetic word problems. Surveying the learning science literature, we posit that the problem-solving process can be split into three distinct steps: text comprehension, solution planning and solution execution. We construct tests for each one in order to understand whether current LLMs display the same cognitive biases as children in these steps. We generate a novel set of word problems for each of these tests, using a neuro-symbolic approach that enables fine-grained control over the problem features. We find evidence that LLMs, with and without instruction-tuning, exhibit human-like biases in both the text-comprehension and the solution-planning steps of the solving process, but not in the final step, in which the arithmetic expressions are executed to obtain the answer.","url_abs":"https://arxiv.org/abs/2401.18070v2","url_pdf":"https://arxiv.org/pdf/2401.18070v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"do-language-models-exhibit-the-same-cognitive","repo_url":"https://github.com/eth-lre/solving-biases","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok"}}],"tasks":[{"task_slug":"reading-comprehension","task_name":"Reading Comprehension"}],"methods":[{"method_slug":"set","method_name":"SET"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=2401.18070","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2401.18070"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/eth-lre/solving-biases","reach":{"status":"ok"}}],"summary":{"ran":9,"ran_violates":1,"unverified":2},"by_repo_kind":{"official":{"samples":12,"ran":10,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":12,"samples":[{"code_sha256_prefix":"86d6ac0a87dfe9d7","entry":"get_entity_n_for_agent","repo":"eth-lre/solving-biases","repo_kind":"official","path":"concate_temp/instantiate_lf.py","file_url":"https://github.com/eth-lre/solving-biases/blob/HEAD/concate_temp/instantiate_lf.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"86d6ac0a87dfe9d7"}},{"code_sha256_prefix":"bebe3264264cd87d","entry":"get_lf_exemplars","repo":"eth-lre/solving-biases","repo_kind":"official","path":"utils/data_utils.py","file_url":"https://github.com/eth-lre/solving-biases/blob/HEAD/utils/data_utils.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"bebe3264264cd87d"}},{"code_sha256_prefix":"a1fb5348c1336d77","entry":"get_special_relations_dict","repo":"eth-lre/solving-biases","repo_kind":"official","path":"concate_temp/instantiate_lf.py","file_url":"https://github.com/eth-lre/solving-biases/blob/HEAD/concate_temp/instantiate_lf.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"a1fb5348c1336d77"}},{"code_sha256_prefix":"9055a78f95bce1d1","entry":"get_subject_from_q","repo":"eth-lre/solving-biases","repo_kind":"official","path":"concate_temp/concate_temp.py","file_url":"https://github.com/eth-lre/solving-biases/blob/HEAD/concate_temp/concate_temp.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"9055a78f95bce1d1"}},{"code_sha256_prefix":"dd69c2e2e8f79e3a","entry":"isFloat","repo":"eth-lre/solving-biases","repo_kind":"official","path":"concate_temp/lf_temp_prepare.py","file_url":"https://github.com/eth-lre/solving-biases/blob/HEAD/concate_temp/lf_temp_prepare.py","link_basis":"plan_row","language":"python","status":"ran_violates","verification_level":1,"contract_check":"VIOLATES","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"dd69c2e2e8f79e3a"}},{"code_sha256_prefix":"b0f1b23d62512994","entry":"is_int","repo":"eth-lre/solving-biases","repo_kind":"official","path":"utils/data_utils.py","file_url":"https://github.com/eth-lre/solving-biases/blob/HEAD/utils/data_utils.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"b0f1b23d62512994"}},{"code_sha256_prefix":"b21bae29bcdb0be7","entry":"load_txt","repo":"eth-lre/solving-biases","repo_kind":"official","path":"utils/data_utils.py","file_url":"https://github.com/eth-lre/solving-biases/blob/HEAD/utils/data_utils.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"b21bae29bcdb0be7"}},{"code_sha256_prefix":"b819792910bc36fe","entry":"load_world_model","repo":"eth-lre/solving-biases","repo_kind":"official","path":"concate_temp/concate_temp.py","file_url":"https://github.com/eth-lre/solving-biases/blob/HEAD/concate_temp/concate_temp.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"b819792910bc36fe"}},{"code_sha256_prefix":"8a7d134df7a067d8","entry":"prompt_funct","repo":"eth-lre/solving-biases","repo_kind":"official","path":"concate_temp/paraphrase.py","file_url":"https://github.com/eth-lre/solving-biases/blob/HEAD/concate_temp/paraphrase.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"8a7d134df7a067d8"}},{"code_sha256_prefix":"fb2fb70dbea4d501","entry":"to_lf_dict","repo":"eth-lre/solving-biases","repo_kind":"official","path":"concate_temp/lf_temp_prepare.py","file_url":"https://github.com/eth-lre/solving-biases/blob/HEAD/concate_temp/lf_temp_prepare.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"fb2fb70dbea4d501"}},{"code_sha256_prefix":"a6a989c10f94cfba","entry":"get_answer_for_lf","repo":"eth-lre/solving-biases","repo_kind":"official","path":"concate_temp/instantiate_lf.py","file_url":"https://github.com/eth-lre/solving-biases/blob/HEAD/concate_temp/instantiate_lf.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"a6a989c10f94cfba"}},{"code_sha256_prefix":"da18088c1461761b","entry":"load_temp_dict","repo":"eth-lre/solving-biases","repo_kind":"official","path":"concate_temp/concate_temp.py","file_url":"https://github.com/eth-lre/solving-biases/blob/HEAD/concate_temp/concate_temp.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"da18088c1461761b"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}