{"url":"/dataset/vist","name":"VIST","full_name":"Visual Storytelling","description_markdown":"The **Visual Storytelling** Dataset (**VIST**) consists of 210,819 unique photos and 50,000 stories. The images were collected from albums on Flickr. The albums included 10 to 50 images and all the images in an album are taken in a 48-hour span. The stories were created by workers on Amazon Mechanical Turk, where the workers were instructed to choose five images from the album and write a story about them. Every story has five sentences, and every sentence is paired with its appropriate image. The dataset is split into 3 subsets, a training set (80%), a validation set (10%) and a test set (10%). All the words and interpunction signs in the stories are separated by a space character and all the location names are replaced with the word location. All the names of people are replaced with the words male or female depending on the gender of the person.\r\n\r\nSource: [Stories for Images-in-Sequence by using Visual and Narrative Components This research was partially funded by Pendulibrium and the Faculty of computer science and engineering, Ss. Cyril and Methodius University in Skopje.](https://arxiv.org/abs/1805.05622)\r\nImage Source: [https://arxiv.org/pdf/1604.03968.pdf](https://arxiv.org/pdf/1604.03968.pdf)","description_withheld":null,"homepage":"https://visionandlanguage.net/VIST/dataset.html","introduced_date":"2013-01-01","introduced_date_note":null,"introduced_by":{"paper":"/paper/visual-storytelling","title":"Visual Storytelling","first_author":"Ting-Hao","url":null},"license":{"name":"Unknown","url":null},"modalities":[{"name":"Images","url":"/datasets/modality/images"},{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Visual Storytelling","url":"/task/visual-storytelling","datasets_with_task":"/datasets/task/visual-storytelling"},{"name":"Story Continuation","url":"/task/story-continuation","datasets_with_task":"/datasets/task/story-continuation"}],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["VIST"],"data_loaders":[{"repo":"https://github.com/visualstorytelling/vist","url":"https://github.com/visualstorytelling/vist","frameworks":[]}],"num_papers_in_archive":107,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/visual-storytelling-on-vist","task":"Visual Storytelling","dataset_variant":"VIST","rows":33,"metrics":["BLEU-4","CIDEr","METEOR","BLEU-1","BLEU-2","BLEU-3","ROUGE-L","SPICE","BLEURT","MLTD"],"first_row_in_archive_order":{"model":"HEGR","paper":"/paper/two-heads-are-better-than-one-hypergraph","metrics":{"BLEU-4":"16.7","CIDEr":"14.1","METEOR":"37.8"},"code_links":[]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/story-continuation-on-vist","task":"Story Continuation","dataset_variant":"VIST","rows":2,"metrics":["FID"],"first_row_in_archive_order":{"model":"AR-LDM (SIS captions)","paper":"/paper/synthesizing-coherent-story-with-auto","metrics":{"FID":"16.95"},"code_links":[{"title":"xichenpan/ARLDM","url":"https://github.com/xichenpan/ARLDM"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[{"paper":"/paper/aog-lstm-an-adaptive-attention-neural-network","title":"AOG-LSTM: An adaptive attention neural network for visual storytelling","date":"2023-06-26","rows_on_this_dataset":1,"code_links":0,"syntology":null},{"paper":"/paper/synthesizing-coherent-story-with-auto","title":"Synthesizing Coherent Story with Auto-Regressive Latent Diffusion Models","date":"2022-11-20","rows_on_this_dataset":2,"code_links":1,"syntology":null},{"paper":"/paper/vision-transformer-based-model-for-describing","title":"Vision Transformer Based Model for Describing a Set of Images as a Story","date":"2022-10-06","rows_on_this_dataset":1,"code_links":0,"syntology":null},{"paper":"/paper/coherent-visual-storytelling-via-parallel-top","title":"Coherent Visual Storytelling via Parallel Top-Down Visual and Topic Attention","date":"2022-08-17","rows_on_this_dataset":1,"code_links":0,"syntology":null},{"paper":"/paper/sentistory-a-multi-layered-sentiment-aware","title":"SentiStory: A Multi-Layered Sentiment-Aware Generative Model for Visual Storytelling","date":"2022-06-16","rows_on_this_dataset":1,"code_links":0,"syntology":null},{"paper":"/paper/visual-storytelling-with-hierarchical-bert","title":"Visual Storytelling with Hierarchical BERT Semantic Guidance","date":"2022-01-10","rows_on_this_dataset":1,"code_links":0,"syntology":null},{"paper":"/paper/two-heads-are-better-than-one-hypergraph","title":"Two Heads are Better Than One: Hypergraph-Enhanced Graph Reasoning for Visual Event Ratiocination","date":"2021-07-18","rows_on_this_dataset":1,"code_links":0,"syntology":null},{"paper":"/paper/transitional-adaptation-of-pretrained-models","title":"Transitional Adaptation of Pretrained Models for Visual Storytelling","date":"2021-06-19","rows_on_this_dataset":2,"code_links":0,"syntology":null},{"paper":"/paper/imagine-reason-and-write-visual-storytelling","title":"Imagine, Reason and Write: Visual Storytelling with Graph Knowledge and Relational Reasoning","date":"2021-05-18","rows_on_this_dataset":1,"code_links":0,"syntology":null},{"paper":"/paper/plot-and-rework-modeling-storylines-for","title":"Plot and Rework: Modeling Storylines for Visual Storytelling","date":"2021-05-14","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/commonsense-knowledge-aware-concept-selection","title":"Commonsense Knowledge Aware Concept Selection For Diverse and Informative Visual Storytelling","date":"2021-02-05","rows_on_this_dataset":1,"code_links":0,"syntology":null},{"paper":"/paper/bert-hlstms-bert-and-hierarchical-lstms-for","title":"BERT-hLSTMs: BERT and Hierarchical LSTMs for Visual Storytelling","date":"2020-12-03","rows_on_this_dataset":2,"code_links":0,"syntology":null},{"paper":"/paper/diverse-and-relevant-visual-storytelling-with","title":"Diverse and Relevant Visual Storytelling with Scene Graph Embeddings","date":"2020-11-01","rows_on_this_dataset":1,"code_links":0,"syntology":null},{"paper":"/paper/hierarchical-memory-decoder-for-visual","title":"Hierarchical memory decoder for visual narrating","date":"2020-09-01","rows_on_this_dataset":1,"code_links":0,"syntology":null},{"paper":"/paper/storytelling-from-an-image-stream-using-scene","title":"Storytelling from an Image Stream Using Scene Graphs","date":"2020-04-03","rows_on_this_dataset":1,"code_links":0,"syntology":null},{"paper":"/paper/hide-and-tell-learning-to-bridge-photo","title":"Hide-and-Tell: Learning to Bridge Photo Streams for Visual Storytelling","date":"2020-02-03","rows_on_this_dataset":1,"code_links":0,"syntology":null},{"paper":"/paper/visual-storytelling-via-predicting-anchor","title":"Visual Storytelling via Predicting Anchor Word Embeddings in the Stories","date":"2020-01-13","rows_on_this_dataset":1,"code_links":0,"syntology":null},{"paper":"/paper/keep-it-consistent-topic-aware-storytelling","title":"Keep it Consistent: Topic-Aware Storytelling from an Image Stream via Iterative Multi-agent Communication","date":"2019-11-11","rows_on_this_dataset":1,"code_links":0,"syntology":null},{"paper":"/paper/what-makes-a-good-story-designing-composite","title":"What Makes A Good Story? Designing Composite Rewards for Visual Storytelling","date":"2019-09-11","rows_on_this_dataset":5,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":12,"samples_ran":0,"samples_unverified":12,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/informative-visual-storytelling-with-cross","title":"Informative Visual Storytelling with Cross-modal Rules","date":"2019-07-07","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/knowledgeable-storyteller-a-commonsense","title":"Knowledgeable Storyteller: A Commonsense-Driven Generative Model for Visual Storytelling","date":"2019-05-04","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/contextualize-show-and-tell-a-neural-visual","title":"Contextualize, Show and Tell: A Neural Visual Storyteller","date":"2018-06-03","rows_on_this_dataset":1,"code_links":2,"syntology":null},{"paper":"/paper/glac-net-glocal-attention-cascading-networks","title":"GLAC Net: GLocal Attention Cascading Networks for Multi-image Cued Story Generation","date":"2018-05-28","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/hierarchically-structured-reinforcement","title":"Hierarchically Structured Reinforcement Learning for Topically Coherent Visual Story Generation","date":"2018-05-21","rows_on_this_dataset":1,"code_links":0,"syntology":null},{"paper":"/paper/no-metrics-are-perfect-adversarial-reward","title":"No Metrics Are Perfect: Adversarial Reward Learning for Visual Storytelling","date":"2018-04-24","rows_on_this_dataset":3,"code_links":2,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":8,"samples_ran":1,"samples_unverified":7,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/hierarchically-attentive-rnn-for-album","title":"Hierarchically-Attentive RNN for Album Summarization and Storytelling","date":"2017-08-09","rows_on_this_dataset":1,"code_links":0,"syntology":null}],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":2,"samples_harvested":20,"samples_ran":1,"samples_unverified":19,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":1,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}