{"url":"/dataset/cold-causal-reasoning-in-closed-daily","name":"COLD: Causal Reasoning in Closed Daily Activities","full_name":null,"description_markdown":"The causal reasoning dataset is generated using the Causal Reasoning in Closed Daily Activities (COLD) framework that helps evaluate large language models (LLMs) on their causal reasoning abilities within real-world, everyday activities. This dataset provides causal questions that simulate common activities such as shopping, baking a cake, riding a bus, planting a tree, and going on a train ride. With approximately 9 million causal queries, the COLD dataset challenges LLMs to understand and reason about the causal relationships between events that are familiar and grounded in human experience.\r\n\r\nEach query consists of a premise (an event) and a pair of choices representing possible causal effects. The goal of the model is to correctly identify which choice is the most plausible cause/effect of the given premise, testing the model's understanding of cause-and-effect relationships.\r\n\r\nKey Features:\r\nActivity Types: The dataset covers various everyday activities: shopping, cake baking, train ride, tree planting, and bus ride.\r\nCausal Queries: Each query includes a premise and two possible causal events (choices). The model must decide which of the two choices is the more likely cause or effect.\r\nMultiple-Choice Format: The queries can be formatted as multiple-choice questions (MCQA), where the model must choose between two options.\r\n\r\nThe dataset provides a valuable test for causal reasoning in NLP models, focusing on realistic, daily-life scenarios.","description_withheld":null,"homepage":"https://huggingface.co/datasets/Exploration-Lab/COLD","introduced_date":"2024-11-29","introduced_date_note":null,"introduced_by":{"paper":"/paper/cold-causal-reasoning-in-closed-daily","title":"COLD: Causal reasOning in cLosed Daily activities","first_author":"Abhinav Joshi","url":null},"license":null,"modalities":[{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Commonsense Causal Reasoning","url":"/task/commonsense-causal-reasoning","datasets_with_task":"/datasets/task/commonsense-causal-reasoning"},{"name":"Event Causality Identification","url":"/task/event-causality-identification","datasets_with_task":"/datasets/task/event-causality-identification"}],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["COLD: Causal Reasoning in Closed Daily Activities"],"data_loaders":[{"repo":"https://github.com/Exploration-Lab/COLD","url":"https://huggingface.co/datasets/Exploration-Lab/COLD","frameworks":[]}],"num_papers_in_archive":1,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-25T09:33:49+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}