Datasets › open-compass/CriticBench
open-compass/CriticBench
[Dataset on HF] [Project Page] [Subjective LeaderBoard] [Objective LeaderBoard]
CriticBench is a novel benchmark designed to comprehensively and reliably evaluate the critique abilities of Large Language Models (LLMs). These critique abilities are crucial for scalable oversight and self-improvement of LLMs. While many recent studies explore how LLMs can judge and refine flaws in their generated outputs, the measurement of critique abilities remains under-explored.
Here are the key aspects of CriticBench:
-
Purpose: To assess LLMs' critique abilities across four dimensions:
- Feedback: How well an LLM provides constructive feedback.
- Comparison: The ability to compare and contrast different responses.
- Refinement: How effectively an LLM can refine flawed or suboptimal outputs.
- Meta-feedback: The LLM's ability to reflect on its own performance.
-
Tasks: CriticBench encompasses nine diverse tasks, each evaluating LLMs' critique abilities at varying levels of quality granularity.
-
Evaluation: The benchmark evaluates both open-source and closed-source LLMs, revealing intriguing relationships between critique abilities, response qualities, and model scales.
-
Resources: Datasets, resources, and an evaluation toolkit for CriticBench will be publicly released.
In summary, CriticBench aims to provide a comprehensive framework for assessing LLMs' critique and self-improvement capabilities, contributing to the advancement of large-scale language models in various applications.
Benchmarks archive 2025-07-28
No leaderboard in the archive resolves to this dataset.
Papers archive 2025-07-28
No paper in the archive has a leaderboard row on this dataset; the archive counts 1 paper for it but never published that list.
Dataset loaders archive 2025-07-28
No loader listed in the archive.
Tasks archive 2025-07-28
No task tagged in the archive.
License archive 2025-07-28
Modalities archive 2025-07-28
Languages archive 2025-07-28
Variants archive 2025-07-28
- open-compass/CriticBench
1 variant name, as the archive lists them.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections