{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/contextualize-show-and-tell-a-neural-visual","title":"Contextualize, Show and Tell: A Neural Visual Storyteller","arxiv_id":"1806.00738","date":"2018-06-03","proceeding":null,"authors":["Diana Gonzalez-Rico","Gibran Fuentes-Pineda"],"abstract":"We present a neural model for generating short stories from image sequences,\nwhich extends the image description model by Vinyals et al. (Vinyals et al.,\n2015). This extension relies on an encoder LSTM to compute a context vector of\neach story from the image sequence. This context vector is used as the first\nstate of multiple independent decoder LSTMs, each of which generates the\nportion of the story corresponding to each image in the sequence by taking the\nimage embedding as the first input. Our model showed competitive results with\nthe METEOR metric and human ratings in the internal track of the Visual\nStorytelling Challenge 2018.","url_abs":"http://arxiv.org/abs/1806.00738v1","url_pdf":"http://arxiv.org/pdf/1806.00738v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"contextualize-show-and-tell-a-neural-visual","repo_url":"https://github.com/dianaglzrico/neural-visual-storyteller","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok","spdx":"Apache-2.0"}},{"paper_slug":"contextualize-show-and-tell-a-neural-visual","repo_url":"https://github.com/dgonzalez-ri/neural-visual-storyteller","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok","spdx":"Apache-2.0"}}],"tasks":[{"task_slug":"decoder","task_name":"Decoder"},{"task_slug":null,"task_name":"Image Description"},{"task_slug":"visual-storytelling","task_name":"Visual Storytelling"}],"methods":[{"method_slug":"lstm","method_name":"LSTM"},{"method_slug":"sigmoid-activation","method_name":"Sigmoid Activation"},{"method_slug":"tanh-activation","method_name":"Tanh Activation"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/visual-storytelling-on-vist","task":"Visual Storytelling","dataset":"VIST","model":"CST","rank_in_archive_order":22,"of":33,"metrics":{"BLEU-1":"60.1","BLEU-2":"36.5","BLEU-3":"21.1","BLEU-4":"12.7","CIDEr":"5.1","METEOR":"34.4","ROUGE-L":"29.2"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1806.00738","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}