{"url":"/dataset/logieval","name":"LogiEval","full_name":null,"description_markdown":"The **LogiEval dataset** is a benchmark suite designed for evaluating the **logical reasoning abilities** of prompt-based language models, particularly instruct-prompt large language models. Here are some key details about LogiEval:\r\n\r\n1. **Purpose and Origin**:\r\n   - LogiEval was created to assess how well language models perform in tasks that require logical reasoning.\r\n   - It is based on the **OpenAI Eval library** and focuses on evaluating logical reasoning abilities.\r\n   - The dataset was developed by researchers to address the need for robust logical reasoning evaluation.\r\n\r\n2. **Contents**:\r\n   - LogiEval contains a set of **logical reasoning tasks** that challenge models to reason deductively.\r\n   - The tasks cover various types of logical reasoning, providing a comprehensive evaluation.\r\n   - The dataset includes **8,678 QA instances** sourced from expert-written questions.\r\n\r\n3. **Usage**:\r\n   - Researchers and practitioners can use LogiEval to assess the logical reasoning capabilities of different models.\r\n   - To utilize LogiEval, one can follow the instructions provided in the repository, including setting up the necessary environment and running evaluations.\r\n\r\n4. **Citation**:\r\n   - If you're interested in using LogiEval or referring to it in your work, you can cite the following paper:\r\n     - **Title**: \"Evaluating the Logical Reasoning Ability of ChatGPT and GPT-4\"\r\n     - **Authors**: Hanmeng Liu, Ruoxi Ning, Zhiyang Teng, Jian Liu, Qiji Zhou, Yue Zhang\r\n     - **Year**: 2023\r\n     - **Link**: [Read the paper](https://arxiv.org/abs/2304.03439) ⁴\r\n\r\nIn summary, LogiEval provides a valuable resource for assessing logical reasoning abilities in prompt-based language models. Researchers can use it to evaluate and compare different models' performance in logical reasoning tasks.\r\n\r\nSource: Conversation with Bing, 3/18/2024\r\n(1) Evaluating the Logical Reasoning Ability of ChatGPT and GPT-4 - arXiv.org. https://arxiv.org/pdf/2304.03439.pdf.\r\n(2) GitHub - csitfun/LogiEval: a benchmark suite for testing logical .... https://github.com/csitfun/LogiEval.\r\n(3) [2007.08124] LogiQA: A Challenge Dataset for Machine Reading .... https://arxiv.org/abs/2007.08124.\r\n(4) [2203.15099] LogicInference: A New Dataset for Teaching Logical .... https://arxiv.org/abs/2203.15099.\r\n(5) Evaluating the Logical Reasoning Ability of ChatGPT and GPT-4. https://arxiv.org/abs/2304.03439.","description_withheld":null,"homepage":"https://github.com/csitfun/LogiEval","introduced_date":"2023-04-07","introduced_date_note":null,"introduced_by":{"paper":"/paper/evaluating-the-logical-reasoning-ability-of","title":"Evaluating the Logical Reasoning Ability of ChatGPT and GPT-4","first_author":"Hanmeng Liu","url":null},"license":null,"modalities":[],"tasks":[],"languages":[],"variants":["LogiEval"],"data_loaders":[],"num_papers_in_archive":2,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}