Browse State-of-the-Art › Visual Storytelling
Visual Storytelling
37 papers with code · 1 benchmark · 4 datasets archive 2025-07-28
( Image credit: No Metrics Are Perfect )
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
1 leaderboard table shown for this task, 1 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| VIST (33 rows) | HEGR | Two Heads are Better Than One: Hypergraph-Enhanced Graph Reasoning... | — | — | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
4 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
1 subtask in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
30 shown of 37 papers with code (115 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
27 Aug 2024 2 repositories listedDesigns and artworks are ubiquitous across various creative fields, requiring graphic design skills and dedicated software to create compositions that include many graphical elements, such as logos, icons, symbols, and…
-
1 Jan 2021 2 repositories listedVisual storytelling and story comprehension are uniquely human skills that play a central role in how we learn about and experience the world.
-
3 Jun 2018 2 repositories listedWe present a neural model for generating short stories from image sequences, which extends the image description model by Vinyals et al.
-
24 Apr 2018 2 repositories listed Syntology ran 1 of 8 samples · 7 unverifiedThough impressive results have been achieved in visual captioning, the task of generating abstract stories from photo streams is still a little-tapped problem.
-
11 Jun 2025 1 repository listedText-to-image generation models have made significant progress in producing high-quality images from textual descriptions, yet they continue to struggle with maintaining subject consistency across multiple images, a…
-
StoryReasoning Dataset: Using Chain-of-Thought for Scene Understanding and Grounded Story Generation15 May 2025 1 repository listedVisual storytelling systems struggle to maintain character identity across frames and link actions to appropriate subjects, frequently leading to referential hallucinations.
-
16 Apr 2025 1 repository listedOver the past years, advances in artificial intelligence (AI) have demonstrated how AI can solve many perception and generation tasks, such as image classification and text writing, yet reasoning remains a challenge.
-
16 Nov 2024 1 repository listedSketch animations offer a powerful medium for visual storytelling, from simple flip-book doodles to professional studio productions.
-
8 Oct 2024 1 repository listedThis tutorial explores recent advancements in multimodal pretrained and large models, capable of integrating and processing diverse data forms such as text, images, audio, and video.
-
7 Aug 2024 1 repository listedRecent image generation models excel at creating high-quality images from brief captions.
-
13 Jul 2024 1 repository listedVisual storytelling involves generating a sequence of coherent frames from a textual storyline while maintaining consistency in characters and scenes.
-
5 Jul 2024 1 repository listedWe then use this method to evaluate the stories generated by several models, showing that the foundation model LLaVA obtains the best result, but only slightly so compared to TAPM, a 50-times smaller visual storytelling…
-
15 Jun 2024 1 repository listedTo address this gap, we introduce CoMM, a high-quality Coherent interleaved image-text MultiModal dataset designed to enhance the coherence, consistency, and alignment of generated multimodal content.
-
22 Apr 2024 1 repository listedContemporary makeup transfer methods primarily focus on replicating makeup from one face to another, considerably limiting their use in creating diverse and creative character makeup essential for visual storytelling.
-
1 Jan 2024 1 repository listedGenerative models have recently exhibited exceptional capabilities in text-to-image generation but still struggle to generate image sequences coherently.
-
3 Nov 2023 1 repository listedYet, the desire to experience manga in vibrant colors has sparked the pursuit of manga colorization, a task of paramount significance for artists.
-
26 Oct 2023 1 repository listed Syntology ran 2 of 3 samples · 1 unverified · 3 pointer-only (licence)A proper evaluation of stories generated for a sequence of images -- the task commonly referred to as visual storytelling -- must consider multiple aspects, such as coherence, grammatical correctness, and visual…
-
6 Oct 2023 1 repository listedIn this paper, we collect an anthology of 100 visual stories from authors who participated in our systematic creative process of improvised story-building based on image sequences.
-
31 Aug 2023 1 repository listedLarge vision-language models (LVLMs) have recently witnessed rapid advancements, exhibiting a remarkable capacity for perceiving, understanding, and processing visual information by connecting visual receptor with large…
-
13 Jul 2023 1 repository listedFor the first module, we leverage an off-the-shelf video retrieval system and extract video depths as motion structure.
-
1 Jun 2023 1 repository listed Syntology ran 0 of 3 samples · 3 unverifiedGenerative models have recently exhibited exceptional capabilities in text-to-image generation, but still struggle to generate image sequences coherently.
-
3 May 2023 1 repository listedHumans can naturally reason from superficial state differences (e.
-
30 Mar 2023 1 repository listedCharacters are essential to the plot of any story.
-
20 Mar 2023 1 repository listedWe present Positional Diffusion, a plug-and-play graph formulation with Diffusion Probabilistic Models to address positional reasoning.
-
1 Jul 2022 1 repository listedWe measure the reliability of our metric sets by analysing its correlation with human judgement scores on a sample of machine stories obtained from 4 state-of-the-arts models trained on the Visual Storytelling Dataset…
-
31 May 2022 1 repository listedThese results depict the effectiveness of commonsense knowledge infusion in improving the performance and expressiveness of scene graph generation for visual understanding and reasoning tasks.
-
8 May 2022 1 repository listedWe measure the reliability of our metric sets by analysing its correlation with human judgement scores on a sample of machine stories obtained from 4 state-of-the-arts models trained on the Visual Storytelling Dataset…
-
1 May 2022 1 repository listedIn this paper, we present the VHED (VIST Human Evaluation Data) dataset, which first re-purposes human evaluation results for automatic evaluation; hence we develop Vrank (VIST Ranker), a novel reference-free VIST…
-
14 May 2021 1 repository listedWriting a coherent and engaging story is not easy.
-
3 Dec 2019 1 repository listedThis paper introduces KG-Story, a three-stage framework that allows the story generation model to take advantage of external Knowledge Graphs to produce interesting stories.
Syntology lines on 3 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections