{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/what-are-the-odds-language-models-are-capable","title":"What Are the Odds? Language Models Are Capable of Probabilistic Reasoning","arxiv_id":"2406.12830","date":"2024-06-18","proceeding":null,"authors":["Akshay Paruchuri","Jake Garrison","Shun Liao","John Hernandez","Jacob Sunshine","Tim Althoff","Xin Liu","Daniel McDuff"],"abstract":"Language models (LM) are capable of remarkably complex linguistic tasks; however, numerical reasoning is an area in which they frequently struggle. An important but rarely evaluated form of reasoning is understanding probability distributions. In this paper, we focus on evaluating the probabilistic reasoning capabilities of LMs using idealized and real-world statistical distributions. We perform a systematic evaluation of state-of-the-art LMs on three tasks: estimating percentiles, drawing samples, and calculating probabilities. We evaluate three ways to provide context to LMs 1) anchoring examples from within a distribution or family of distributions, 2) real-world context, 3) summary statistics on which to base a Normal approximation. Models can make inferences about distributions, and can be further aided by the incorporation of real-world context, example shots and simplified assumptions, even if these assumptions are incorrect or misspecified. To conduct this work, we developed a comprehensive benchmark distribution dataset with associated question-answer pairs that we have released publicly.","url_abs":"https://arxiv.org/abs/2406.12830v3","url_pdf":"https://arxiv.org/pdf/2406.12830v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"what-are-the-odds-language-models-are-capable","repo_url":"https://github.com/yahskapar/LLMs-and-Probabilistic-Reasoning","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[],"methods":[{"method_slug":"base","method_name":"BASE"},{"method_slug":"focus","method_name":"Focus"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/2406.12830","atlas_url":"https://app.syntology.ai/?focus=2406.12830","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2406.12830"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-25T09:33:49+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/yahskapar/LLMs-and-Probabilistic-Reasoning","reach":null}],"summary":{"ran_fixture":3},"by_repo_kind":{"official":{"samples":3,"ran":3,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"dad56d7c94ef462c","entry":"calculate_probability_within_range","repo":"yahskapar/LLMs-and-Probabilistic-Reasoning","repo_kind":"official","path":"generation/idealized_generation/idealized_distributions.py","file_url":"https://github.com/yahskapar/LLMs-and-Probabilistic-Reasoning/blob/HEAD/generation/idealized_generation/idealized_distributions.py","link_basis":"first_harvest_node","language":"python","status":"ran_fixture","verification_level":1,"contract_check":"RAISES","metamorphic_tier":"invariant","behaviour_fingerprint":true,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"dad56d7c94ef462c"}},{"code_sha256_prefix":"e0d260ba5eaa4e22","entry":"calculate_target_percentile_values","repo":"yahskapar/LLMs-and-Probabilistic-Reasoning","repo_kind":"official","path":"generation/idealized_generation/idealized_distributions.py","file_url":"https://github.com/yahskapar/LLMs-and-Probabilistic-Reasoning/blob/HEAD/generation/idealized_generation/idealized_distributions.py","link_basis":"first_harvest_node","language":"python","status":"ran_fixture","verification_level":1,"contract_check":"RAISES","metamorphic_tier":"deterministic","behaviour_fingerprint":true,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"e0d260ba5eaa4e22"}},{"code_sha256_prefix":"ca7c8b6cf5b2e442","entry":"calculate_target_ranges","repo":"yahskapar/LLMs-and-Probabilistic-Reasoning","repo_kind":"official","path":"generation/idealized_generation/idealized_distributions.py","file_url":"https://github.com/yahskapar/LLMs-and-Probabilistic-Reasoning/blob/HEAD/generation/idealized_generation/idealized_distributions.py","link_basis":"first_harvest_node","language":"python","status":"ran_fixture","verification_level":1,"contract_check":"RAISES","metamorphic_tier":"deterministic","behaviour_fingerprint":true,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"ca7c8b6cf5b2e442"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}