{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/is-chatgpt-a-good-causal-reasoner-a","title":"Is ChatGPT a Good Causal Reasoner? A Comprehensive Evaluation","arxiv_id":"2305.07375","date":"2023-05-12","proceeding":null,"authors":["Jinglong Gao","Xiao Ding","Bing Qin","Ting Liu"],"abstract":"Causal reasoning ability is crucial for numerous NLP applications. Despite the impressive emerging ability of ChatGPT in various NLP tasks, it is unclear how well ChatGPT performs in causal reasoning. In this paper, we conduct the first comprehensive evaluation of the ChatGPT's causal reasoning capabilities. Experiments show that ChatGPT is not a good causal reasoner, but a good causal explainer. Besides, ChatGPT has a serious hallucination on causal reasoning, possibly due to the reporting biases between causal and non-causal relationships in natural language, as well as ChatGPT's upgrading processes, such as RLHF. The In-Context Learning (ICL) and Chain-of-Thought (CoT) techniques can further exacerbate such causal hallucination. Additionally, the causal reasoning ability of ChatGPT is sensitive to the words used to express the causal concept in prompts, and close-ended prompts perform better than open-ended prompts. For events in sentences, ChatGPT excels at capturing explicit causality rather than implicit causality, and performs better in sentences with lower event density and smaller lexical distance between events. The code is available on https://github.com/ArrogantL/ChatGPT4CausalReasoning .","url_abs":"https://arxiv.org/abs/2305.07375v4","url_pdf":"https://arxiv.org/pdf/2305.07375v4.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"is-chatgpt-a-good-causal-reasoner-a","repo_url":"https://github.com/ArrogantL/ChatGPT4CausalReasoning","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"Apache-2.0"}}],"tasks":[{"task_slug":"hallucination","task_name":"Hallucination"},{"task_slug":"in-context-learning","task_name":"In-Context Learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/2305.07375","atlas_url":"https://app.syntology.ai/?focus=2305.07375","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2305.07375"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-25T09:33:49+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/ArrogantL/ChatGPT4CausalReasoning","reach":{"status":"ok","spdx":"Apache-2.0"}}],"summary":{"unverified":3},"by_repo_kind":{"official":{"samples":3,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"8967aefe849dc141","entry":"completion_with_backoff_chat","repo":"ArrogantL/ChatGPT4CausalReasoning","repo_kind":"official","path":"CD_binary_classification.py","file_url":"https://github.com/ArrogantL/ChatGPT4CausalReasoning/blob/HEAD/CD_binary_classification.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"8967aefe849dc141"}},{"code_sha256_prefix":"d977e31ed8c92307","entry":"completion_with_backoff_davinci","repo":"ArrogantL/ChatGPT4CausalReasoning","repo_kind":"official","path":"CD_binary_classification.py","file_url":"https://github.com/ArrogantL/ChatGPT4CausalReasoning/blob/HEAD/CD_binary_classification.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"d977e31ed8c92307"}},{"code_sha256_prefix":"3706ba5f4c89273f","entry":"eval_task1_in_pos_neg","repo":"ArrogantL/ChatGPT4CausalReasoning","repo_kind":"official","path":"CD_and_CEG_compute_score.py","file_url":"https://github.com/ArrogantL/ChatGPT4CausalReasoning/blob/HEAD/CD_and_CEG_compute_score.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"3706ba5f4c89273f"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}