{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/vgbench-evaluating-large-language-models-on","title":"VGBench: Evaluating Large Language Models on Vector Graphics Understanding and Generation","arxiv_id":"2407.10972","date":"2024-07-15","proceeding":null,"authors":["Bocheng Zou","Mu Cai","Jianrui Zhang","Yong Jae Lee"],"abstract":"In the realm of vision models, the primary mode of representation is using pixels to rasterize the visual world. Yet this is not always the best or unique way to represent visual content, especially for designers and artists who depict the world using geometry primitives such as polygons. Vector graphics (VG), on the other hand, offer a textual representation of visual content, which can be more concise and powerful for content like cartoons, sketches and scientific figures. Recent studies have shown promising results on processing vector graphics with capable Large Language Models (LLMs). However, such works focus solely on qualitative results, understanding, or a specific type of vector graphics. We propose VGBench, a comprehensive benchmark for LLMs on handling vector graphics through diverse aspects, including (a) both visual understanding and generation, (b) evaluation of various vector graphics formats, (c) diverse question types, (d) wide range of prompting techniques, (e) under multiple LLMs and (f) comparison with VLMs on rasterized representations. Evaluating on our collected 4279 understanding and 5845 generation samples, we find that LLMs show strong capability on both aspects while exhibiting less desirable performance on low-level formats (SVG). Both data and evaluation pipeline will be open-sourced at https://vgbench.github.io.","url_abs":"https://arxiv.org/abs/2407.10972v2","url_pdf":"https://arxiv.org/pdf/2407.10972v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"vgbench-evaluating-large-language-models-on","repo_url":"https://github.com/vgbench/VGBench","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok","spdx":"Apache-2.0"}}],"tasks":[{"task_slug":"vector-graphics","task_name":"Vector Graphics"}],"methods":[{"method_slug":"focus","method_name":"Focus"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2407.10972","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2407.10972"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/vgbench/VGBench","reach":{"status":"ok","spdx":"Apache-2.0"}}],"summary":{"ran":3,"unverified":3},"by_repo_kind":{"official":{"samples":6,"ran":3,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"e87679a61667de51","entry":"convert_sample_format","repo":"vgbench/VGBench","repo_kind":"official","path":"evaluate.py","file_url":"https://github.com/vgbench/VGBench/blob/HEAD/evaluate.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"e87679a61667de51"}},{"code_sha256_prefix":"f377c500498620f8","entry":"generate_system_message","repo":"vgbench/VGBench","repo_kind":"official","path":"generate_questions.py","file_url":"https://github.com/vgbench/VGBench/blob/HEAD/generate_questions.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"f377c500498620f8"}},{"code_sha256_prefix":"d54d6d247c2fb8dc","entry":"load_questions","repo":"vgbench/VGBench","repo_kind":"official","path":"build_dataset.py","file_url":"https://github.com/vgbench/VGBench/blob/HEAD/build_dataset.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"d54d6d247c2fb8dc"}},{"code_sha256_prefix":"ffd8c42a422834f1","entry":"calculate_activation_statistics","repo":"vgbench/VGBench","repo_kind":"official","path":"evaluate_fid_score.py","file_url":"https://github.com/vgbench/VGBench/blob/HEAD/evaluate_fid_score.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"ffd8c42a422834f1"}},{"code_sha256_prefix":"0def50a351111624","entry":"calculate_frechet_distance","repo":"vgbench/VGBench","repo_kind":"official","path":"evaluate_fid_score.py","file_url":"https://github.com/vgbench/VGBench/blob/HEAD/evaluate_fid_score.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":"RAISES","metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"0def50a351111624"}},{"code_sha256_prefix":"3f1d9512c08824f8","entry":"get_activations","repo":"vgbench/VGBench","repo_kind":"official","path":"evaluate_fid_score.py","file_url":"https://github.com/vgbench/VGBench/blob/HEAD/evaluate_fid_score.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"3f1d9512c08824f8"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}