Papers › Keep it Consistent: Topic-Aware Storytelling from an Image Stream via Iterative...
Keep it Consistent: Topic-Aware Storytelling from an Image Stream via Iterative Multi-agent Communication
Ruize Wang, Zhongyu Wei, Ying Cheng, Piji Li, Haijun Shan, Ji Zhang, Qi Zhang, Xuanjing Huang
Visual storytelling aims to generate a narrative paragraph from a sequence of images automatically. Existing approaches construct text description independently for each image and roughly concatenate them as a story, which leads to the problem of generating semantically incoherent content. In this paper, we propose a new way for visual storytelling by introducing a topic description task to detect the global semantic context of an image stream. A story is then constructed with the guidance of the topic description. In order to combine the two generation tasks, we propose a multi-agent communication framework that regards the topic description generator and the story generator as two agents and learn them simultaneously via iterative updating mechanism. We validate our approach on VIST dataset, where quantitative results, ablations, and human evaluation demonstrate our method's good ability in generating stories with higher quality compared to state-of-the-art methods.
In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.
Code
No code repository is listed for this paper in the archive or in Syntology's graph.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| Visual Storytelling | VIST | TAVST (RL) | BLEU-1 | 64.2 | #9 of 33 | Archive leaderboard | report |
| Visual Storytelling | VIST | TAVST (RL) | BLEU-2 | 39.6 | #9 of 33 | Archive leaderboard | report |
| Visual Storytelling | VIST | TAVST (RL) | BLEU-3 | 23.7 | #9 of 33 | Archive leaderboard | report |
| Visual Storytelling | VIST | TAVST (RL) | BLEU-4 | 14.6 | #9 of 33 | Archive leaderboard | report |
| Visual Storytelling | VIST | TAVST (RL) | CIDEr | 9.2 | #9 of 33 | Archive leaderboard | report |
| Visual Storytelling | VIST | TAVST (RL) | METEOR | 35.7 | #9 of 33 | Archive leaderboard | report |
| Visual Storytelling | VIST | TAVST (RL) | ROUGE-L | 31 | #9 of 33 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections