Datasets › LogiEval

LogiEval

Introduced by Hanmeng Liu et al. in Evaluating the Logical Reasoning Ability of ChatGPT and GPT-47 Apr 2023 archive 2025-07-28

The LogiEval dataset is a benchmark suite designed for evaluating the logical reasoning abilities of prompt-based language models, particularly instruct-prompt large language models. Here are some key details about LogiEval:

  1. Purpose and Origin:
  2. LogiEval was created to assess how well language models perform in tasks that require logical reasoning.
  3. It is based on the OpenAI Eval library and focuses on evaluating logical reasoning abilities.
  4. The dataset was developed by researchers to address the need for robust logical reasoning evaluation.

  5. Contents:

  6. LogiEval contains a set of logical reasoning tasks that challenge models to reason deductively.
  7. The tasks cover various types of logical reasoning, providing a comprehensive evaluation.
  8. The dataset includes 8,678 QA instances sourced from expert-written questions.

  9. Usage:

  10. Researchers and practitioners can use LogiEval to assess the logical reasoning capabilities of different models.
  11. To utilize LogiEval, one can follow the instructions provided in the repository, including setting up the necessary environment and running evaluations.

  12. Citation:

  13. If you're interested in using LogiEval or referring to it in your work, you can cite the following paper:
    • Title: "Evaluating the Logical Reasoning Ability of ChatGPT and GPT-4"
    • Authors: Hanmeng Liu, Ruoxi Ning, Zhiyang Teng, Jian Liu, Qiji Zhou, Yue Zhang
    • Year: 2023
    • Link: Read the paper ⁴

In summary, LogiEval provides a valuable resource for assessing logical reasoning abilities in prompt-based language models. Researchers can use it to evaluate and compare different models' performance in logical reasoning tasks.

Source: Conversation with Bing, 3/18/2024 (1) Evaluating the Logical Reasoning Ability of ChatGPT and GPT-4 - arXiv.org. https://arxiv.org/pdf/2304.03439.pdf. (2) GitHub - csitfun/LogiEval: a benchmark suite for testing logical .... https://github.com/csitfun/LogiEval. (3) [2007.08124] LogiQA: A Challenge Dataset for Machine Reading .... https://arxiv.org/abs/2007.08124. (4) [2203.15099] LogicInference: A New Dataset for Teaching Logical .... https://arxiv.org/abs/2203.15099. (5) Evaluating the Logical Reasoning Ability of ChatGPT and GPT-4. https://arxiv.org/abs/2304.03439.

Benchmarks archive 2025-07-28

No leaderboard in the archive resolves to this dataset.

Papers archive 2025-07-28

No paper in the archive has a leaderboard row on this dataset; the archive counts 2 papers for it but never published that list.

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

No task tagged in the archive.

License archive 2025-07-28

No licence recorded in the archive. Absence here is not a statement about the dataset's terms.

Modalities archive 2025-07-28

No modality tagged.

Languages archive 2025-07-28

No language tagged.

Variants archive 2025-07-28

  • LogiEval

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections