Papers › Knowledgeable Storyteller: A Commonsense-Driven Generative Model for Visual Storytelling

Knowledgeable Storyteller: A Commonsense-Driven Generative Model for Visual Storytelling

4 May 2019IJCAI 2019 2019 5archive 2025-07-28

Pengcheng Yang, Fuli Luo, Peng Chen, Lei LI, Zhiyi Yin, Xiaodong He, Xu sun

The visual storytelling (VST) task aims at generating a reasonable and coherent paragraph-level story with the image stream as input. Different from caption that is a direct and literal description of image content, the story in the VST task tends to contain plenty of imaginary concepts that do not appear in the image. This requires the AI agent to reason and associate with the imaginary concepts based on implicit commonsense knowledge to generate a reasonable story describing the image stream. Therefore, in this work, we present a commonsensedriven generative model, which aims to introduce crucial commonsense from the external knowledge base for visual storytelling. Our approach first extracts a set of candidate knowledge graphs from the knowledge base. Then, an elaborately designed vision-aware directional encoding schema is adopted to effectively integrate the most informative commonsense. Besides, we strive to maximize the semantic similarity within the output during decoding to enhance the coherence of the generated text. Results show that our approach can outperform the state-of-the-art systems by a large margin, which achieves a 29% relative improvement of CIDEr score. With additional commonsense and semantic-relevance based objective, the generated stories are more diverse and coherent.

PaperPDFCode

Code

lancopku/CVST mentioned in paper report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

AI AgentKnowledge GraphsSemantic SimilaritySemantic Textual SimilarityText GenerationVisual Storytelling

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Visual Storytelling VIST K-Storyteller BLEU-4 12.8 #21 of 33 Archive leaderboard report
Visual Storytelling VIST K-Storyteller CIDEr 12.1 #21 of 33 Archive leaderboard report
Visual Storytelling VIST K-Storyteller METEOR 35.2 #21 of 33 Archive leaderboard report
Visual Storytelling VIST K-Storyteller ROUGE-L 29.9 #21 of 33 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections