{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/mmt-image-guided-story-ending-generation-with","title":"MMT: Image-guided Story Ending Generation with Multimodal Memory Transformer","arxiv_id":null,"date":"2022-10-10","proceeding":"ACM MM 2022 10","authors":["Dizhan Xue","Shengsheng Qian","Quan Fang","Changsheng Xu"],"abstract":"As a specific form of story generation, Image-guided Story Ending Generation (IgSEG) is a recently proposed task of generating a story ending for a given multi-sentence story plot and an ending-related image. Unlike existing image captioning tasks or story ending generation tasks, IgSEG aims to generate a factual description that conforms to both the contextual logic and the relevant visual concepts. To date, existing methods for IgSEG ignore the relationships between the multimodal information and do not integrate multimodal features appropriately. Therefore, in this work, we propose Multimodal Memory Transformer (MMT), an end-to-end framework that models and fuses both contextual and visual information to effectively capture the multimodal dependency for IgSEG. Firstly, we extract textual and visual features separately by employing modality-specific large-scale pretrained encoders. Secondly, we utilize the memory-augmented cross-modal attention network to learn cross-modal relationships and conduct the fine-grained feature fusion effectively. Finally, a multimodal transformer decoder constructs attention among multimodal features to learn the story dependency and generates informative, reasonable, and coherent story endings. In experiments, extensive automatic evaluation results and human evaluation results indicate the significant performance boost of our proposed MMT over state-of-the-art methods on two benchmark datasets.","url_abs":"https://dl.acm.org/doi/abs/10.1145/3503161.3548022","url_pdf":"https://dl.acm.org/doi/pdf/10.1145/3503161.3548022","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"mmt-image-guided-story-ending-generation-with","repo_url":"https://github.com/LivXue/MMT","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"decoder","task_name":"Decoder"},{"task_slug":"image-captioning","task_name":"Image Captioning"},{"task_slug":"image-guided-story-ending-generation","task_name":"Image-guided Story Ending Generation"},{"task_slug":"sentence","task_name":"Sentence"},{"task_slug":"story-generation","task_name":"Story Generation"}],"methods":[{"method_slug":"absolute-position-encodings","method_name":"Absolute Position Encodings"},{"method_slug":"adam","method_name":"Adam"},{"method_slug":"attention","method_name":"Attention"},{"method_slug":"bpe","method_name":"BPE"},{"method_slug":"dense-connections","method_name":"Dense Connections"},{"method_slug":"dropout","method_name":"Dropout"},{"method_slug":"label-smoothing","method_name":"Label Smoothing"},{"method_slug":"layer-normalization","method_name":"Layer Normalization"},{"method_slug":"linear-layer","method_name":"Linear Layer"},{"method_slug":"multi-head-attention","method_name":"Multi-Head Attention"},{"method_slug":"position-wise-feed-forward-layer","method_name":"Position-Wise Feed-Forward Layer"},{"method_slug":"residual-connection","method_name":"Residual Connection"},{"method_slug":"softmax","method_name":"Softmax"},{"method_slug":"transformer","method_name":"Transformer"}],"datasets_introduced":[{"slug":"lsmdc-e","name":"LSMDC-E","full_name":"LSMDC-Ending"}],"methods_introduced":[],"results":[{"leaderboard":"/sota/image-guided-story-ending-generation-on-lsmdc","task":"Image-guided Story Ending Generation","dataset":"LSMDC-E","model":"MMT","rank_in_archive_order":1,"of":4,"metrics":{"BLEU-1":"18.52","BLEU-2":"5.99","BLEU-3":"2.51","BLEU-4":"1.13","CIDEr":"12.41","METEOR":"12.87","ROUGE-L":"20.99"},"uses_additional_data":false},{"leaderboard":"/sota/image-guided-story-ending-generation-on-vist","task":"Image-guided Story Ending Generation","dataset":"VIST-E","model":"MMT","rank_in_archive_order":1,"of":6,"metrics":{"BLEU-1":"22.87","BLEU-2":"8.68","BLEU-3":"4.38","BLEU-4":"2.61","CIDEr":"25.41","METEOR":"15.55","ROUGE-L":"23.61"},"uses_additional_data":false}],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}