{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/hierarchical-photo-scene-encoder-for-album","title":"Hierarchical Photo-Scene Encoder for Album Storytelling","arxiv_id":"1902.00669","date":"2019-02-02","proceeding":null,"authors":["Bairui Wang","Lin Ma","Wei zhang","Wenhao Jiang","Feng Zhang"],"abstract":"In this paper, we propose a novel model with a hierarchical photo-scene\nencoder and a reconstructor for the task of album storytelling. The photo-scene\nencoder contains two sub-encoders, namely the photo and scene encoders, which\nare stacked together and behave hierarchically to fully exploit the structure\ninformation of the photos within an album. Specifically, the photo encoder\ngenerates semantic representation for each photo while exploiting temporal\nrelationships among them. The scene encoder, relying on the obtained photo\nrepresentations, is responsible for detecting the scene changes and generating\nscene representations. Subsequently, the decoder dynamically and attentively\nsummarizes the encoded photo and scene representations to generate a sequence\nof album representations, based on which a story consisting of multiple\ncoherent sentences is generated. In order to fully extract the useful semantic\ninformation from an album, a reconstructor is employed to reproduce the\nsummarized album representations based on the hidden states of the decoder. The\nproposed model can be trained in an end-to-end manner, which results in an\nimproved performance over the state-of-the-arts on the public visual\nstorytelling (VIST) dataset. Ablation studies further demonstrate the\neffectiveness of the proposed hierarchical photo-scene encoder and\nreconstructor.","url_abs":"http://arxiv.org/abs/1902.00669v1","url_pdf":"http://arxiv.org/pdf/1902.00669v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"decoder","task_name":"Decoder"},{"task_slug":"image-guided-story-ending-generation","task_name":"Image-guided Story Ending Generation"},{"task_slug":"visual-storytelling","task_name":"Visual Storytelling"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/image-guided-story-ending-generation-on-vist","task":"Image-guided Story Ending Generation","dataset":"VIST-E","model":"T-CVAE","rank_in_archive_order":5,"of":6,"metrics":{"BLEU-1":"14.34","BLEU-2":"5.06","BLEU-3":"2.01","BLEU-4":"1.13","CIDEr":"11.49","METEOR":"4.23","ROUGE-L":"15.51"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/1902.00669","atlas_url":"https://app.syntology.ai/?focus=1902.00669","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}