Papers › Commonsense Knowledge Aware Concept Selection For Diverse and Informative Visual Storytelling
Commonsense Knowledge Aware Concept Selection For Diverse and Informative Visual Storytelling
Hong Chen, Yifei HUANG, Hiroya Takamura, Hideki Nakayama
Visual storytelling is a task of generating relevant and interesting stories for given image sequences. In this work we aim at increasing the diversity of the generated stories while preserving the informative content from the images. We propose to foster the diversity and informativeness of a generated story by using a concept selection module that suggests a set of concept candidates. Then, we utilize a large scale pre-trained model to convert concepts and images into full stories. To enrich the candidate concepts, a commonsense knowledge graph is created for each image sequence from which the concept candidates are proposed. To obtain appropriate concepts from the graph, we propose two novel modules that consider the correlation among candidate concepts and the image-concept correlation. Extensive automatic and human evaluation results demonstrate that our model can produce reasonable concepts. This enables our model to outperform the previous models by a large margin on the diversity and informativeness of the story, while retaining the relevance of the story to the image sequence.
In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.
Code
No code repository is listed for this paper in the archive or in Syntology's graph.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| Visual Storytelling | VIST | MCSM+RNN | BLEU-3 | 23.1 | #19 of 33 | Archive leaderboard | report |
| Visual Storytelling | VIST | MCSM+RNN | BLEU-4 | 13 | #19 of 33 | Archive leaderboard | report |
| Visual Storytelling | VIST | MCSM+RNN | CIDEr | 11 | #19 of 33 | Archive leaderboard | report |
| Visual Storytelling | VIST | MCSM+RNN | METEOR | 36.1 | #19 of 33 | Archive leaderboard | report |
| Visual Storytelling | VIST | MCSM+RNN | ROUGE-L | 30.7 | #19 of 33 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections