{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/incorporating-global-visual-features-into","title":"Incorporating Global Visual Features into Attention-Based Neural Machine Translation","arxiv_id":"1701.06521","date":"2017-01-23","proceeding":null,"authors":["Iacer Calixto","Qun Liu","Nick Campbell"],"abstract":"We introduce multi-modal, attention-based neural machine translation (NMT)\nmodels which incorporate visual features into different parts of both the\nencoder and the decoder. We utilise global image features extracted using a\npre-trained convolutional neural network and incorporate them (i) as words in\nthe source sentence, (ii) to initialise the encoder hidden state, and (iii) as\nadditional data to initialise the decoder hidden state. In our experiments, we\nevaluate how these different strategies to incorporate global image features\ncompare and which ones perform best. We also study the impact that adding\nsynthetic multi-modal, multilingual data brings and find that the additional\ndata have a positive impact on multi-modal models. We report new\nstate-of-the-art results and our best models also significantly improve on a\ncomparable phrase-based Statistical MT (PBSMT) model trained on the Multi30k\ndata set according to all metrics evaluated. To the best of our knowledge, it\nis the first time a purely neural model significantly improves over a PBSMT\nmodel on all metrics evaluated on this data set.","url_abs":"http://arxiv.org/abs/1701.06521v1","url_pdf":"http://arxiv.org/pdf/1701.06521v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"decoder","task_name":"Decoder"},{"task_slug":"machine-translation","task_name":"Machine Translation"},{"task_slug":"multimodal-machine-translation","task_name":"Multimodal Machine Translation"},{"task_slug":"nmt","task_name":"NMT"},{"task_slug":"sentence","task_name":"Sentence"},{"task_slug":"translation","task_name":"Translation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/multimodal-machine-translation-on-multi30k","task":"Multimodal Machine Translation","dataset":"Multi30K","model":"IMGD","rank_in_archive_order":10,"of":15,"metrics":{"BLEU (EN-DE)":"37.3","Meteor (EN-DE)":"55.1"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1701.06521","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}