Papers › Visual Storytelling via Predicting Anchor Word Embeddings in the Stories

Visual Storytelling via Predicting Anchor Word Embeddings in the Stories

13 Jan 2020arXiv:2001.04541archive 2025-07-28

Bowen Zhang, Hexiang Hu, Fei Sha

We propose a learning model for the task of visual storytelling. The main idea is to predict anchor word embeddings from the images and use the embeddings and the image features jointly to generate narrative sentences. We use the embeddings of randomly sampled nouns from the groundtruth stories as the target anchor word embeddings to learn the predictor. To narrate a sequence of images, we use the predicted anchor word embeddings and the image features as the joint input to a seq2seq model. As opposed to state-of-the-art methods, the proposed model is simple in design, easy to optimize, and attains the best results in most automatic evaluation metrics. In human evaluation, the method also outperforms competing methods.

PaperPDF

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Visual StorytellingWord Embeddings

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Visual Storytelling VIST StoryAnchor: w/ Predicted Nouns BLEU-1 65.1 #15 of 33 Archive leaderboard report
Visual Storytelling VIST StoryAnchor: w/ Predicted Nouns BLEU-2 40.0 #15 of 33 Archive leaderboard report
Visual Storytelling VIST StoryAnchor: w/ Predicted Nouns BLEU-3 23.4 #15 of 33 Archive leaderboard report
Visual Storytelling VIST StoryAnchor: w/ Predicted Nouns BLEU-4 14 #15 of 33 Archive leaderboard report
Visual Storytelling VIST StoryAnchor: w/ Predicted Nouns CIDEr 9.9 #15 of 33 Archive leaderboard report
Visual Storytelling VIST StoryAnchor: w/ Predicted Nouns METEOR 35.5 #15 of 33 Archive leaderboard report
Visual Storytelling VIST StoryAnchor: w/ Predicted Nouns ROUGE-L 30 #15 of 33 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

LSTMSeq2SeqSigmoid ActivationTanh Activation

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections