{"url":"/dataset/criticbench","name":"CriticBench","full_name":null,"description_markdown":"**CriticBench** is a comprehensive benchmark designed to assess the abilities of **Large Language Models (LLMs)** to critique and rectify their reasoning across various tasks. It encompasses five reasoning domains:\r\n\r\n1. **Mathematical**\r\n2. **Commonsense**\r\n3. **Symbolic**\r\n4. **Coding**\r\n5. **Algorithmic**\r\n\r\nCriticBench compiles **15 datasets** and incorporates responses from **three LLM families**. By utilizing CriticBench, researchers evaluate and dissect the performance of **17 LLMs** in **generation**, **critique**, and **correction reasoning** (referred to as **GQC reasoning**). Notable findings include:\r\n\r\n1. A **linear relationship** in GQC capabilities, with critique-focused training significantly enhancing performance.\r\n2. **Task-dependent variation** in correction effectiveness, with logic-oriented tasks being more amenable to correction.\r\n3. **GQC knowledge inconsistencies** that decrease as model size increases.\r\n4. An intriguing **inter-model critiquing dynamic**, where stronger models excel at critiquing weaker ones, while weaker models surprisingly surpass stronger ones in self-critique.\r\n\r\n(1) CriticBench: Benchmarking LLMs for Critique-Correct Reasoning. https://arxiv.org/abs/2402.14809.\r\n(2) CriticBench: Benchmarking LLMs for Critique-Correct Reasoning. http://export.arxiv.org/abs/2402.14809.\r\n(3) CriticBench: Benchmarking LLMs for Critique-Correct Reasoning. https://openreview.net/forum?id=sc5i7q6DQO.\r\n(4) CriticBench: Benchmarking LLMs for Critique-Correct Reasoning - arXiv.org. https://arxiv.org/html/2402.14809v2.\r\n(5) undefined. https://doi.org/10.48550/arXiv.2402.14809.","description_withheld":null,"homepage":"https://criticbench.github.io/","introduced_date":"2024-02-22","introduced_date_note":null,"introduced_by":{"paper":"/paper/criticbench-benchmarking-llms-for-critique","title":"CriticBench: Benchmarking LLMs for Critique-Correct Reasoning","first_author":"Zicheng Lin","url":null},"license":{"name":"MIT","url":null},"modalities":[{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Code Generation","url":"/task/code-generation","datasets_with_task":"/datasets/task/code-generation"},{"name":"Common Sense Reasoning","url":"/task/common-sense-reasoning","datasets_with_task":"/datasets/task/common-sense-reasoning"},{"name":"Mathematical Reasoning","url":"/task/mathematical-reasoning","datasets_with_task":"/datasets/task/mathematical-reasoning"},{"name":"Code Repair","url":"/task/code-repair","datasets_with_task":"/datasets/task/code-repair"}],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["CriticBench"],"data_loaders":[{"repo":"https://github.com/huggingface/datasets","url":"https://huggingface.co/datasets/llm-agents/CriticBench","frameworks":["tf","pytorch","jax"]}],"num_papers_in_archive":10,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}