Papers › Hide-and-Tell: Learning to Bridge Photo Streams for Visual Storytelling

Hide-and-Tell: Learning to Bridge Photo Streams for Visual Storytelling

3 Feb 2020arXiv:2002.00774archive 2025-07-28

Yunjae Jung, Dahun Kim, Sanghyun Woo, Kyung-Su Kim, Sungjin Kim, In So Kweon

Visual storytelling is a task of creating a short story based on photo streams. Unlike existing visual captioning, storytelling aims to contain not only factual descriptions, but also human-like narration and semantics. However, the VIST dataset consists only of a small, fixed number of photos per story. Therefore, the main challenge of visual storytelling is to fill in the visual gap between photos with narrative and imaginative story. In this paper, we propose to explicitly learn to imagine a storyline that bridges the visual gap. During training, one or more photos is randomly omitted from the input stack, and we train the network to produce a full plausible story even with missing photo(s). Furthermore, we propose for visual storytelling a hide-and-tell model, which is designed to learn non-local relations across the photo streams and to refine and improve conventional RNN-based models. In experiments, we show that our scheme of hide-and-tell, and the network design are indeed effective at storytelling, and that our model outperforms previous state-of-the-art methods in automatic metrics. Finally, we qualitatively show the learned ability to interpolate storyline over visual gaps.

PaperPDF

In Syntology View this paper on Syntology: its repositories, every harvested function with whether it ran, its licence and the call to fetch it.

Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Image CaptioningVisual Storytelling

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Visual Storytelling VIST INet BLEU-1 64.4 #7 of 33 Archive leaderboard report
Visual Storytelling VIST INet BLEU-2 0.401 #7 of 33 Archive leaderboard report
Visual Storytelling VIST INet BLEU-3 23.9 #7 of 33 Archive leaderboard report
Visual Storytelling VIST INet BLEU-4 14.7 #7 of 33 Archive leaderboard report
Visual Storytelling VIST INet CIDEr 10 #7 of 33 Archive leaderboard report
Visual Storytelling VIST INet METEOR 35.6 #7 of 33 Archive leaderboard report
Visual Storytelling VIST INet ROUGE-L 29.7 #7 of 33 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections