{"url":"/dataset/2d-atoms","name":"2D-ATOMS","full_name":null,"description_markdown":"Official dataset for **Towards A Holistic Landscape of Situated Theory of Mind in Large Language Models.** *Ziqiao Ma, Jacob Sansom, Run Peng, Joyce Chai. EMNLP Findings, 2023.*\r\n\r\nWe introduce **2D-ATOMS** dataset, a novel text-based dataset that evaluates a machine's reasoning process under a situated theory-of-mind setting.\r\n\r\nOur dataset includes 9 different ToM evaluation tasks for each mental state under ATOMS framework, and 1 reality-checking task to test LLMs’ understanding of the world. It is important to acknowledge that our experiment serves as a proof of concept and does not aim to cover the entire spectrum of machine ToM, as our case studies are far from being exhaustive or systematic. Here we release the zero-shot version of our dataset, which is used in our paper.","description_withheld":null,"homepage":"https://huggingface.co/datasets/sled-umich/2D-ATOMS","introduced_date":"2023-10-30","introduced_date_note":null,"introduced_by":{"paper":"/paper/towards-a-holistic-landscape-of-situated","title":"Towards A Holistic Landscape of Situated Theory of Mind in Large Language Models","first_author":"Ziqiao Ma","url":null},"license":{"name":"MIT","url":"https://huggingface.co/datasets/sled-umich/2D-ATOMS"},"modalities":[{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Theory of Mind Modeling","url":"/task/theory-of-mind-modeling","datasets_with_task":"/datasets/task/theory-of-mind-modeling"}],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["2D-ATOMS"],"data_loaders":[],"num_papers_in_archive":1,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}