{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/humaneval-v-evaluating-visual-understanding","title":"HumanEval-V: Evaluating Visual Understanding and Reasoning Abilities of Large Multimodal Models Through Coding Tasks","arxiv_id":"2410.12381","date":"2024-10-16","proceeding":null,"authors":["Fengji Zhang","Linquan Wu","Huiyu Bai","Guancheng Lin","Xiao Li","Xiao Yu","Yue Wang","Bei Chen","Jacky Keung"],"abstract":"Coding tasks have been valuable for evaluating Large Language Models (LLMs), as they demand the comprehension of high-level instructions, complex reasoning, and the implementation of functional programs -- core capabilities for advancing Artificial General Intelligence. Despite the progress in Large Multimodal Models (LMMs), which extend LLMs with visual perception and understanding capabilities, there remains a notable lack of coding benchmarks that rigorously assess these models, particularly in tasks that emphasize visual reasoning. To address this gap, we introduce HumanEval-V, a novel and lightweight benchmark specifically designed to evaluate LMMs' visual understanding and reasoning capabilities through code generation. HumanEval-V includes 108 carefully crafted, entry-level Python coding tasks derived from platforms like CodeForces and Stack Overflow. Each task is adapted by modifying the context and algorithmic patterns of the original problems, with visual elements redrawn to ensure distinction from the source, preventing potential data leakage. LMMs are required to complete the code solution based on the provided visual context and a predefined Python function signature outlining the task requirements. Every task is equipped with meticulously handcrafted test cases to ensure a thorough and reliable evaluation of model-generated solutions. We evaluate 19 state-of-the-art LMMs using HumanEval-V, uncovering significant challenges. Proprietary models like GPT-4o achieve only 13% pass@1 and 36.4% pass@10, while open-weight models with 70B parameters score below 4% pass@1. Ablation studies further reveal the limitations of current LMMs in vision reasoning and coding capabilities. These results underscore key areas for future research to enhance LMMs' capabilities. We have open-sourced our code and benchmark at https://github.com/HumanEval-V/HumanEval-V-Benchmark.","url_abs":"https://arxiv.org/abs/2410.12381v2","url_pdf":"https://arxiv.org/pdf/2410.12381v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"humaneval-v-evaluating-visual-understanding","repo_url":"https://github.com/HumanEval-V/HumanEval-V-Benchmark","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok"}}],"tasks":[{"task_slug":"code-generation","task_name":"Code Generation"},{"task_slug":"humaneval","task_name":"HumanEval"},{"task_slug":"visual-reasoning","task_name":"Visual Reasoning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2410.12381","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2410.12381"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/HumanEval-V/HumanEval-V-Benchmark","reach":{"status":"ok"}}],"summary":{"ran":6},"by_repo_kind":{"official":{"samples":6,"ran":6,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":6,"samples":[{"code_sha256_prefix":"f0c3af44403d4d56","entry":"check_correctness","repo":"HumanEval-V/HumanEval-V-Benchmark","repo_kind":"official","path":"execution.py","file_url":"https://github.com/HumanEval-V/HumanEval-V-Benchmark/blob/HEAD/execution.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"f0c3af44403d4d56"}},{"code_sha256_prefix":"8ef282a84b503ff7","entry":"extract_code_without_function_class_parent","repo":"HumanEval-V/HumanEval-V-Benchmark","repo_kind":"official","path":"evaluate.py","file_url":"https://github.com/HumanEval-V/HumanEval-V-Benchmark/blob/HEAD/evaluate.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"8ef282a84b503ff7"}},{"code_sha256_prefix":"4ce220ab82541813","entry":"load_and_append_json","repo":"HumanEval-V/HumanEval-V-Benchmark","repo_kind":"official","path":"utils.py","file_url":"https://github.com/HumanEval-V/HumanEval-V-Benchmark/blob/HEAD/utils.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"4ce220ab82541813"}},{"code_sha256_prefix":"2d7add6459ff5b37","entry":"load_json","repo":"HumanEval-V/HumanEval-V-Benchmark","repo_kind":"official","path":"utils.py","file_url":"https://github.com/HumanEval-V/HumanEval-V-Benchmark/blob/HEAD/utils.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"2d7add6459ff5b37"}},{"code_sha256_prefix":"a917ba970738a94f","entry":"post_process","repo":"HumanEval-V/HumanEval-V-Benchmark","repo_kind":"official","path":"evaluate.py","file_url":"https://github.com/HumanEval-V/HumanEval-V-Benchmark/blob/HEAD/evaluate.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"a917ba970738a94f"}},{"code_sha256_prefix":"c0e3e76bb71a2cc8","entry":"read_file","repo":"HumanEval-V/HumanEval-V-Benchmark","repo_kind":"official","path":"utils.py","file_url":"https://github.com/HumanEval-V/HumanEval-V-Benchmark/blob/HEAD/utils.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"c0e3e76bb71a2cc8"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}