{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/a-visual-attention-grounding-neural-model-for","title":"A Visual Attention Grounding Neural Model for Multimodal Machine Translation","arxiv_id":"1808.08266","date":"2018-08-24","proceeding":"EMNLP 2018 10","authors":["Mingyang Zhou","Runxiang Cheng","Yong Jae Lee","Zhou Yu"],"abstract":"We introduce a novel multimodal machine translation model that utilizes\nparallel visual and textual information. Our model jointly optimizes the\nlearning of a shared visual-language embedding and a translator. The model\nleverages a visual attention grounding mechanism that links the visual\nsemantics with the corresponding textual semantics. Our approach achieves\ncompetitive state-of-the-art results on the Multi30K and the Ambiguous COCO\ndatasets. We also collected a new multilingual multimodal product description\ndataset to simulate a real-world international online shopping scenario. On\nthis dataset, our visual attention grounding model outperforms other methods by\na large margin.","url_abs":"http://arxiv.org/abs/1808.08266v2","url_pdf":"http://arxiv.org/pdf/1808.08266v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"a-visual-attention-grounding-neural-model-for","repo_url":"https://github.com/Eurus-Holmes/VAG-NMT","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"machine-translation","task_name":"Machine Translation"},{"task_slug":"multimodal-machine-translation","task_name":"Multimodal Machine Translation"},{"task_slug":"translation","task_name":"Translation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/multimodal-machine-translation-on-multi30k","task":"Multimodal Machine Translation","dataset":"Multi30K","model":"VAG-NMT","rank_in_archive_order":12,"of":15,"metrics":{"BLEU (EN-DE)":"31.6","Meteor (EN-DE)":"52.2","Meteor (EN-FR)":"70.3"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1808.08266","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}