Datasets › LogiEval
LogiEval
The LogiEval dataset is a benchmark suite designed for evaluating the logical reasoning abilities of prompt-based language models, particularly instruct-prompt large language models. Here are some key details about LogiEval:
- Purpose and Origin:
- LogiEval was created to assess how well language models perform in tasks that require logical reasoning.
- It is based on the OpenAI Eval library and focuses on evaluating logical reasoning abilities.
-
The dataset was developed by researchers to address the need for robust logical reasoning evaluation.
-
Contents:
- LogiEval contains a set of logical reasoning tasks that challenge models to reason deductively.
- The tasks cover various types of logical reasoning, providing a comprehensive evaluation.
-
The dataset includes 8,678 QA instances sourced from expert-written questions.
-
Usage:
- Researchers and practitioners can use LogiEval to assess the logical reasoning capabilities of different models.
-
To utilize LogiEval, one can follow the instructions provided in the repository, including setting up the necessary environment and running evaluations.
-
Citation:
- If you're interested in using LogiEval or referring to it in your work, you can cite the following paper:
- Title: "Evaluating the Logical Reasoning Ability of ChatGPT and GPT-4"
- Authors: Hanmeng Liu, Ruoxi Ning, Zhiyang Teng, Jian Liu, Qiji Zhou, Yue Zhang
- Year: 2023
- Link: Read the paper ⁴
In summary, LogiEval provides a valuable resource for assessing logical reasoning abilities in prompt-based language models. Researchers can use it to evaluate and compare different models' performance in logical reasoning tasks.
Source: Conversation with Bing, 3/18/2024 (1) Evaluating the Logical Reasoning Ability of ChatGPT and GPT-4 - arXiv.org. https://arxiv.org/pdf/2304.03439.pdf. (2) GitHub - csitfun/LogiEval: a benchmark suite for testing logical .... https://github.com/csitfun/LogiEval. (3) [2007.08124] LogiQA: A Challenge Dataset for Machine Reading .... https://arxiv.org/abs/2007.08124. (4) [2203.15099] LogicInference: A New Dataset for Teaching Logical .... https://arxiv.org/abs/2203.15099. (5) Evaluating the Logical Reasoning Ability of ChatGPT and GPT-4. https://arxiv.org/abs/2304.03439.
Benchmarks archive 2025-07-28
No leaderboard in the archive resolves to this dataset.
Papers archive 2025-07-28
No paper in the archive has a leaderboard row on this dataset; the archive counts 2 papers for it but never published that list.
Dataset loaders archive 2025-07-28
No loader listed in the archive.
Tasks archive 2025-07-28
No task tagged in the archive.
License archive 2025-07-28
No licence recorded in the archive. Absence here is not a statement about the dataset's terms.
Modalities archive 2025-07-28
No modality tagged.
Languages archive 2025-07-28
No language tagged.
Variants archive 2025-07-28
- LogiEval
1 variant name, as the archive lists them.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections