{"url":"/dataset/nphardeval","name":"NPHardEval","full_name":null,"description_markdown":"**NPHardEval** is a dynamic benchmark designed to assess the reasoning abilities of **Large Language Models (LLMs)** across a broad spectrum of algorithmic questions. Let's delve into the details:\r\n\r\n1. **Benchmark Purpose**:\r\n   - **Complex Reasoning Ability**: One of the most crucial features of current LLMs is their ability to handle complex reasoning. This capability plays an integral role in complex decision-making tasks.\r\n   - **Inadequacy of Existing Benchmarks**: While several benchmarks exist to evaluate LLMs' reasoning abilities, they fall short in providing a rigorous assessment of the full extent of these abilities. Additionally, publicly accessible and static benchmarks risk overfitting, allowing models to tailor their responses to specific metrics.\r\n   - **Introducing NPHardEval**: To address these limitations, NPHardEval was introduced. It aims to rigorously evaluate LLMs' reasoning abilities by extending up to the **NP-Hard complexity class**.\r\n\r\n2. **Key Features of NPHardEval**:\r\n   - **900 Algorithmic Questions**: NPHardEval includes a diverse set of 900 algorithmic questions, carefully chosen to represent a wide range of complexity classes below NP-Hard. These questions serve as a rigorous measure of LLMs' reasoning abilities.\r\n   - **Dynamic Update Mechanism**: Unlike static benchmarks, NPHardEval dynamically updates its datapoints on a monthly basis. Regular updates mitigate the risk of overfitting, ensuring a more accurate and reliable assessment of LLMs' reasoning capabilities.\r\n\r\n3. **Research Contribution**:\r\n   - **Objective Perspective**: NPHardEval sheds light on the current state of reasoning in LLMs by comparing their performance across complex classes.\r\n   - **Available Resources**: The benchmark dataset and code for NPHardEval are accessible [here](https://arxiv.org/abs/2312.14890) ¹.\r\n\r\nIn summary, NPHardEval provides a comprehensive evaluation framework for assessing LLMs' reasoning abilities through the lens of computational complexity classes. 🌟\r\n\r\n(1) NPHardEval: Dynamic Benchmark on Reasoning Ability of Large Language .... https://arxiv.org/abs/2312.14890.\r\n(2) NPHardEval/README.md at main · casmlab/NPHardEval · GitHub. https://github.com/casmlab/NPHardEval/blob/main/README.md.\r\n(3) NPHardEval: Benchmarking Reasoning Ability of Large Language Models via .... https://frankling2020.github.io/publication/nphardeval/.\r\n(4) undefined. https://doi.org/10.48550/arXiv.2312.14890.","description_withheld":null,"homepage":"https://github.com/casmlab/nphardeval","introduced_date":"2023-12-22","introduced_date_note":null,"introduced_by":{"paper":"/paper/nphardeval-dynamic-benchmark-on-reasoning","title":"NPHardEval: Dynamic Benchmark on Reasoning Ability of Large Language Models via Complexity Classes","first_author":"Lizhou Fan","url":null},"license":null,"modalities":[],"tasks":[],"languages":[],"variants":["NPHardEval"],"data_loaders":[],"num_papers_in_archive":11,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}