{"url":"/dataset/inferential-strategies","name":"Inferential-Strategies","full_name":null,"description_markdown":"A collection of large languge model responses to tasks of propositional logic. The responses are annotated according to the following criteria:\r\n\r\n- Inferential strategy employed by the model. Strategies considered are: *supposition following*, *chain construction*, *compound strategy*, *concatenation strategy*, and the *symbolic strategy*. Binary labels are assigned to each strategy, indicating whether the strategy is present in the model's response.\r\n- Assessment of the validity of the model's final conclusion. Binary labels are assigned to each response, indicating whether the model's conclusion is accurate (\"valid_conclusion\").\r\n- Evaluation of the soundness of the model's rationale. Binary labels are assigned to each response, indicating whether the rationale provided by the model is sound (\"sound_reasoning\").\r\n- A description of the model's reasoning error. This information is provided in the form of a string (\"reasoning_errors\").","description_withheld":null,"homepage":"https://huggingface.co/datasets/mainlp/inferential_strategies","introduced_date":"2024-02-20","introduced_date_note":null,"introduced_by":{"paper":"/paper/comparing-inferential-strategies-of-humans","title":"Comparing Inferential Strategies of Humans and Large Language Models in Deductive Reasoning","first_author":"Philipp Mondorf","url":null},"license":{"name":"cc-by-4.0","url":"https://creativecommons.org/licenses/by-sa/4.0/"},"modalities":[{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Logical Reasoning","url":"/task/logical-reasoning","datasets_with_task":"/datasets/task/logical-reasoning"}],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["Inferential-Strategies"],"data_loaders":[],"num_papers_in_archive":1,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}