{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/autobench-v-can-large-vision-language-models","title":"AutoBench-V: Can Large Vision-Language Models Benchmark Themselves?","arxiv_id":"2410.21259","date":"2024-10-28","proceeding":null,"authors":["Han Bao","Yue Huang","Yanbo Wang","Jiayi Ye","Xiangqi Wang","Xiuying Chen","Yue Zhao","Tianyi Zhou","Mohamed Elhoseiny","Xiangliang Zhang"],"abstract":"Large Vision-Language Models (LVLMs) have become essential for advancing the integration of visual and linguistic information. However, the evaluation of LVLMs presents significant challenges as the evaluation benchmark always demands lots of human cost for its construction, and remains static, lacking flexibility once constructed. Even though automatic evaluation has been explored in textual modality, the visual modality remains under-explored. As a result, in this work, we address a question: \"Can LVLMs themselves be used to benchmark each other in the visual automatically domain?\". We introduce AutoBench-V, an automated framework for serving evaluation on demand, i.e., benchmarking LVLMs based on specific aspects of model capability. AutoBench-V leverages text-to-image models to generate relevant image samples and then utilizes LVLMs to orchestrate visual question-answering (VQA) tasks, completing the evaluation process efficiently and flexibly. Through an extensive evaluation of nine popular LVLMs across five demanded user inputs (i.e., evaluation capabilities), the framework shows effectiveness and reliability.","url_abs":"https://arxiv.org/abs/2410.21259v3","url_pdf":"https://arxiv.org/pdf/2410.21259v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"autobench-v-can-large-vision-language-models","repo_url":"https://github.com/wad3birch/AutoBench-V","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok"}}],"tasks":[{"task_slug":"benchmarking","task_name":"Benchmarking"},{"task_slug":"question-answering","task_name":"Question Answering"},{"task_slug":"visual-question-answering-1","task_name":"Visual Question Answering"},{"task_slug":"visual-question-answering","task_name":"Visual Question Answering (VQA)"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2410.21259","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2410.21259"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/wad3birch/AutoBench-V","reach":{"status":"ok"}}],"summary":{"unverified":4},"by_repo_kind":{"official":{"samples":4,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":4,"samples":[{"code_sha256_prefix":"deeecdc9c4a04e97","entry":"add_corners","repo":"wad3birch/AutoBench-V","repo_kind":"official","path":"process/human_eval.py","file_url":"https://github.com/wad3birch/AutoBench-V/blob/HEAD/process/human_eval.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"deeecdc9c4a04e97"}},{"code_sha256_prefix":"5daf1194bf2db299","entry":"load_config","repo":"wad3birch/AutoBench-V","repo_kind":"official","path":"process/aspect_generate.py","file_url":"https://github.com/wad3birch/AutoBench-V/blob/HEAD/process/aspect_generate.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"5daf1194bf2db299"}},{"code_sha256_prefix":"f2d9761231eeb385","entry":"load_data","repo":"wad3birch/AutoBench-V","repo_kind":"official","path":"process/human_eval.py","file_url":"https://github.com/wad3birch/AutoBench-V/blob/HEAD/process/human_eval.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"f2d9761231eeb385"}},{"code_sha256_prefix":"e670ccbd528c1fa5","entry":"load_json","repo":"wad3birch/AutoBench-V","repo_kind":"official","path":"process/aspect_generate.py","file_url":"https://github.com/wad3birch/AutoBench-V/blob/HEAD/process/aspect_generate.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"e670ccbd528c1fa5"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}