{"url":"/dataset/nphardeval4v","name":"NPHardEval4V","full_name":null,"description_markdown":"**NPHardEval4V** is a dynamic reasoning benchmark designed to evaluate the reasoning capabilities of **Multimodal Large Language Models (MLLMs)**. Let me provide you with more details:\r\n\r\n1. **Purpose and Gap Addressed**:\r\n   - The benchmark aims to address existing gaps in evaluating the pure reasoning abilities of MLLMs.\r\n   - It provides a venue to disentangle the effects of various factors (such as image recognition and instruction following) from the overall performance of the models.\r\n   - By focusing solely on reasoning abilities, NPHardEval4V helps researchers understand and guide further development in this area.\r\n\r\n2. **Construction and Features**:\r\n   - NPHardEval4V is built by converting textual descriptions of questions from the existing NPHardEval dataset into image representations.\r\n   - Unlike traditional benchmarks that primarily focus on static evaluations, NPHardEval4V is dynamic. It is updated monthly to prevent overfitting and ensure authentic and fine-grained model evaluation.\r\n   - The benchmark evaluates MLLMs across three problem classes: polynomial time, NP-complete, and NP-hard problems.\r\n   - It assesses performance in three dimensions:\r\n     - **Recognition (RA)**: Ability to understand image and video modalities.\r\n     - **Instruction-following (ER)**: How well the model follows instructions.\r\n     - **Reasoning (AA)**: Pure reasoning abilities.\r\n\r\n3. **Findings and Impact**:\r\n   - Significant discrepancies in reasoning abilities exist across different models.\r\n   - MLLMs exhibit relatively weak performance compared to Large Language Models (LLMs) in terms of reasoning.\r\n   - Investigating different prompting styles (visual, text, and combined) reveals varying impacts of multimodal inputs on model performance.\r\n\r\nIn summary, NPHardEval4V provides a valuable resource for assessing reasoning abilities in MLLMs and contributes to advancing research in this domain. 🌟","description_withheld":null,"homepage":"https://github.com/lizhouf/nphardeval4v","introduced_date":"2024-03-04","introduced_date_note":null,"introduced_by":{"paper":"/paper/nphardeval4v-a-dynamic-reasoning-benchmark-of","title":"NPHardEval4V: A Dynamic Reasoning Benchmark of Multimodal Large Language Models","first_author":"Lizhou Fan","url":null},"license":null,"modalities":[],"tasks":[],"languages":[],"variants":["NPHardEval4V"],"data_loaders":[],"num_papers_in_archive":4,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}