Papers › Hierarchically Structured Reinforcement Learning for Topically Coherent Visual Story Generation
Hierarchically Structured Reinforcement Learning for Topically Coherent Visual Story Generation
Qiuyuan Huang, Zhe Gan, Asli Celikyilmaz, Dapeng Wu, Jian-Feng Wang, Xiaodong He
We propose a hierarchically structured reinforcement learning approach to address the challenges of planning for generating coherent multi-sentence stories for the visual storytelling task. Within our framework, the task of generating a story given a sequence of images is divided across a two-level hierarchical decoder. The high-level decoder constructs a plan by generating a semantic concept (i.e., topic) for each image in sequence. The low-level decoder generates a sentence for each image using a semantic compositional network, which effectively grounds the sentence generation conditioned on the topic. The two decoders are jointly trained end-to-end using reinforcement learning. We evaluate our model on the visual storytelling (VIST) dataset. Empirical results from both automatic and human evaluations demonstrate that the proposed hierarchically structured reinforced training achieves significantly better performance compared to a strong flat deep reinforcement learning baseline.
In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.
Code
No code repository is listed for this paper in the archive or in Syntology's graph.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| Visual Storytelling | VIST | HSRL w/ Joint Training | BLEU-4 | 12.32 | #24 of 33 | Archive leaderboard | report |
| Visual Storytelling | VIST | HSRL w/ Joint Training | CIDEr | 10.71 | #24 of 33 | Archive leaderboard | report |
| Visual Storytelling | VIST | HSRL w/ Joint Training | METEOR | 35.23 | #24 of 33 | Archive leaderboard | report |
| Visual Storytelling | VIST | HSRL w/ Joint Training | ROUGE-L | 30.84 | #24 of 33 | Archive leaderboard | report |
| Visual Storytelling | VIST | HSRL w/ Joint Training | SPICE | 12.97 | #24 of 33 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections