{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/glac-net-glocal-attention-cascading-networks","title":"GLAC Net: GLocal Attention Cascading Networks for Multi-image Cued Story Generation","arxiv_id":"1805.10973","date":"2018-05-28","proceeding":null,"authors":["Taehyeong Kim","Min-Oh Heo","Seonil Son","Kyoung-Wha Park","Byoung-Tak Zhang"],"abstract":"The task of multi-image cued story generation, such as visual storytelling\ndataset (VIST) challenge, is to compose multiple coherent sentences from a\ngiven sequence of images. The main difficulty is how to generate image-specific\nsentences within the context of overall images. Here we propose a deep learning\nnetwork model, GLAC Net, that generates visual stories by combining\nglobal-local (glocal) attention and context cascading mechanisms. The model\nincorporates two levels of attention, i.e., overall encoding level and image\nfeature level, to construct image-dependent sentences. While standard attention\nconfiguration needs a large number of parameters, the GLAC Net implements them\nin a very simple way via hard connections from the outputs of encoders or image\nfeatures onto the sentence generators. The coherency of the generated story is\nfurther improved by conveying (cascading) the information of the previous\nsentence to the next sentence serially. We evaluate the performance of the GLAC\nNet on the visual storytelling dataset (VIST) and achieve very competitive\nresults compared to the state-of-the-art techniques. Our code and pre-trained\nmodels are available here.","url_abs":"http://arxiv.org/abs/1805.10973v3","url_pdf":"http://arxiv.org/pdf/1805.10973v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"glac-net-glocal-attention-cascading-networks","repo_url":"https://github.com/tkim-snu/GLACNet","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"sentence","task_name":"Sentence"},{"task_slug":"story-generation","task_name":"Story Generation"},{"task_slug":"visual-storytelling","task_name":"Visual Storytelling"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/visual-storytelling-on-vist","task":"Visual Storytelling","dataset":"VIST","model":"GLAC Net","rank_in_archive_order":33,"of":33,"metrics":{"METEOR":"30.14"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1805.10973","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}