Datasets › VIST

VIST (Visual Storytelling)

Introduced by Ting-Hao et al. in Visual Storytelling1 Jan 2013 archive 2025-07-28

The Visual Storytelling Dataset (VIST) consists of 210,819 unique photos and 50,000 stories. The images were collected from albums on Flickr. The albums included 10 to 50 images and all the images in an album are taken in a 48-hour span. The stories were created by workers on Amazon Mechanical Turk, where the workers were instructed to choose five images from the album and write a story about them. Every story has five sentences, and every sentence is paired with its appropriate image. The dataset is split into 3 subsets, a training set (80%), a validation set (10%) and a test set (10%). All the words and interpunction signs in the stories are separated by a space character and all the location names are replaced with the word location. All the names of people are replaced with the words male or female depending on the gender of the person.

Source: Stories for Images-in-Sequence by using Visual and Narrative Components This research was partially funded by Pendulibrium and the Faculty of computer science and engineering, Ss. Cyril and Methodius University in Skopje. Image Source: https://arxiv.org/pdf/1604.03968.pdf

Benchmarks archive 2025-07-28

All 2 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Visual Storytelling VIST HEGR BLEU-4 16.7 Two Heads are Better Than One: Hypergraph-Enhanced Graph... — 33 Compare
Story Continuation VIST AR-LDM (SIS captions) FID 16.95 Synthesizing Coherent Story with Auto-Regressive Latent... xichenpan/ARLDM 2 Compare

Papers archive 2025-07-28

26 shown of 26 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 107. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
AOG-LSTM: An adaptive attention neural network for visual storytelling 0 1 26 Jun 2023 not harvested
Synthesizing Coherent Story with Auto-Regressive Latent Diffusion Models 1 2 20 Nov 2022 not harvested
Vision Transformer Based Model for Describing a Set of Images as a Story 0 1 6 Oct 2022 not harvested
Coherent Visual Storytelling via Parallel Top-Down Visual and Topic Attention 0 1 17 Aug 2022 not harvested
SentiStory: A Multi-Layered Sentiment-Aware Generative Model for Visual Storytelling 0 1 16 Jun 2022 not harvested
Visual Storytelling with Hierarchical BERT Semantic Guidance 0 1 10 Jan 2022 not harvested
Two Heads are Better Than One: Hypergraph-Enhanced Graph Reasoning for Visual Event Ratiocination 0 1 18 Jul 2021 not harvested
Transitional Adaptation of Pretrained Models for Visual Storytelling 0 2 19 Jun 2021 not harvested
Imagine, Reason and Write: Visual Storytelling with Graph Knowledge and Relational Reasoning 0 1 18 May 2021 not harvested
Plot and Rework: Modeling Storylines for Visual Storytelling 1 1 14 May 2021 not harvested
Commonsense Knowledge Aware Concept Selection For Diverse and Informative Visual Storytelling 0 1 5 Feb 2021 not harvested
BERT-hLSTMs: BERT and Hierarchical LSTMs for Visual Storytelling 0 2 3 Dec 2020 not harvested
Diverse and Relevant Visual Storytelling with Scene Graph Embeddings 0 1 1 Nov 2020 not harvested
Hierarchical memory decoder for visual narrating 0 1 1 Sep 2020 not harvested
Storytelling from an Image Stream Using Scene Graphs 0 1 3 Apr 2020 not harvested
Hide-and-Tell: Learning to Bridge Photo Streams for Visual Storytelling 0 1 3 Feb 2020 not harvested
Visual Storytelling via Predicting Anchor Word Embeddings in the Stories 0 1 13 Jan 2020 not harvested
Keep it Consistent: Topic-Aware Storytelling from an Image Stream via Iterative Multi-agent Communication 0 1 11 Nov 2019 not harvested
What Makes A Good Story? Designing Composite Rewards for Visual Storytelling 1 5 11 Sep 2019 ran 0 of 12 samples (12 unverified)
Informative Visual Storytelling with Cross-modal Rules 1 1 7 Jul 2019 not harvested
Knowledgeable Storyteller: A Commonsense-Driven Generative Model for Visual Storytelling 1 1 4 May 2019 not harvested
Contextualize, Show and Tell: A Neural Visual Storyteller 2 1 3 Jun 2018 not harvested
GLAC Net: GLocal Attention Cascading Networks for Multi-image Cued Story Generation 1 1 28 May 2018 not harvested
Hierarchically Structured Reinforcement Learning for Topically Coherent Visual Story Generation 0 1 21 May 2018 not harvested
No Metrics Are Perfect: Adversarial Reward Learning for Visual Storytelling 2 3 24 Apr 2018 ran 1 of 8 samples (7 unverified)
Hierarchically-Attentive RNN for Album Summarization and Storytelling 0 1 9 Aug 2017 not harvested

Dataset loaders archive 2025-07-28

1 loader as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

Unknown

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • VIST

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections