{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/vistorybench-comprehensive-benchmark-suite","title":"ViStoryBench: Comprehensive Benchmark Suite for Story Visualization","arxiv_id":"2505.24862","date":"2025-05-30","proceeding":null,"authors":["Cailin Zhuang","Ailin Huang","Wei Cheng","Jingwei Wu","Yaoqi Hu","Jiaqi Liao","Zhewei Huang","Hongyuan Wang","Xinyao Liao","Weiwei Cai","Hengyuan Xu","Xuanyang Zhang","Xianfang Zeng","Gang Yu","Chi Zhang"],"abstract":"Story visualization, which aims to generate a sequence of visually coherent images aligning with a given narrative and reference images, has seen significant progress with recent advancements in generative models. To further enhance the performance of story visualization frameworks in real-world scenarios, we introduce a comprehensive evaluation benchmark, ViStoryBench. We collect a diverse dataset encompassing various story types and artistic styles, ensuring models are evaluated across multiple dimensions such as different plots (e.g., comedy, horror) and visual aesthetics (e.g., anime, 3D renderings). ViStoryBench is carefully curated to balance narrative structures and visual elements, featuring stories with single and multiple protagonists to test models' ability to maintain character consistency. Additionally, it includes complex plots and intricate world-building to challenge models in generating accurate visuals. To ensure comprehensive comparisons, our benchmark incorporates a wide range of evaluation metrics assessing critical aspects. This structured and multifaceted framework enables researchers to thoroughly identify both the strengths and weaknesses of different models, fostering targeted improvements.","url_abs":"https://arxiv.org/abs/2505.24862v1","url_pdf":"https://arxiv.org/pdf/2505.24862v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"vistorybench-comprehensive-benchmark-suite","repo_url":"https://github.com/vistorybench/vistorybench","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"story-visualization","task_name":"Story Visualization"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2505.24862","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2505.24862"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/vistorybench/vistorybench","reach":null}],"summary":{"ran_draft_wrong":1,"unverified":1},"by_repo_kind":{"official":{"samples":2,"ran":1,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"9f4471a999bc0cf0","entry":"load_config","repo":"vistorybench/vistorybench","repo_kind":"official","path":"vistorybench/dataset_loader/dataset_load.py","file_url":"https://github.com/vistorybench/vistorybench/blob/HEAD/vistorybench/dataset_loader/dataset_load.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":true,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"9f4471a999bc0cf0"}},{"code_sha256_prefix":"d3862ddd961af0e6","entry":"load_story_data","repo":"vistorybench/vistorybench","repo_kind":"official","path":"vistorybench/dataset_loader/dataset_load.py","file_url":"https://github.com/vistorybench/vistorybench/blob/HEAD/vistorybench/dataset_loader/dataset_load.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"d3862ddd961af0e6"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}