Datasets › open-compass/CriticBench

open-compass/CriticBench

Introduced by Tian Lan et al. in CriticEval: Evaluating Large Language Model as Critic21 Feb 2024 archive 2025-07-28

[arXiv] [license]

[Dataset on HF] [Project Page] [Subjective LeaderBoard] [Objective LeaderBoard]

CriticBench is a novel benchmark designed to comprehensively and reliably evaluate the critique abilities of Large Language Models (LLMs). These critique abilities are crucial for scalable oversight and self-improvement of LLMs. While many recent studies explore how LLMs can judge and refine flaws in their generated outputs, the measurement of critique abilities remains under-explored.

Here are the key aspects of CriticBench:

  1. Purpose: To assess LLMs' critique abilities across four dimensions:

    • Feedback: How well an LLM provides constructive feedback.
    • Comparison: The ability to compare and contrast different responses.
    • Refinement: How effectively an LLM can refine flawed or suboptimal outputs.
    • Meta-feedback: The LLM's ability to reflect on its own performance.
  2. Tasks: CriticBench encompasses nine diverse tasks, each evaluating LLMs' critique abilities at varying levels of quality granularity.

  3. Evaluation: The benchmark evaluates both open-source and closed-source LLMs, revealing intriguing relationships between critique abilities, response qualities, and model scales.

  4. Resources: Datasets, resources, and an evaluation toolkit for CriticBench will be publicly released.

In summary, CriticBench aims to provide a comprehensive framework for assessing LLMs' critique and self-improvement capabilities, contributing to the advancement of large-scale language models in various applications.

Benchmarks archive 2025-07-28

No leaderboard in the archive resolves to this dataset.

Papers archive 2025-07-28

No paper in the archive has a leaderboard row on this dataset; the archive counts 1 paper for it but never published that list.

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

No task tagged in the archive.

License archive 2025-07-28

Apache License, Version 2.0

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • open-compass/CriticBench

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections