{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/pervasive-attention-2d-convolutional-neural-1","title":"Pervasive Attention: 2D Convolutional Neural Networks for Sequence-to-Sequence Prediction","arxiv_id":"1808.03867","date":"2018-08-11","proceeding":"CONLL 2018 10","authors":["Maha Elbayad","Laurent Besacier","Jakob Verbeek"],"abstract":"Current state-of-the-art machine translation systems are based on\nencoder-decoder architectures, that first encode the input sequence, and then\ngenerate an output sequence based on the input encoding. Both are interfaced\nwith an attention mechanism that recombines a fixed encoding of the source\ntokens based on the decoder state. We propose an alternative approach which\ninstead relies on a single 2D convolutional neural network across both\nsequences. Each layer of our network re-codes source tokens on the basis of the\noutput sequence produced so far. Attention-like properties are therefore\npervasive throughout the network. Our model yields excellent results,\noutperforming state-of-the-art encoder-decoder systems, while being\nconceptually simpler and having fewer parameters.","url_abs":"http://arxiv.org/abs/1808.03867v3","url_pdf":"http://arxiv.org/pdf/1808.03867v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"pervasive-attention-2d-convolutional-neural-1","repo_url":"https://github.com/elbayadm/attn2d","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"pervasive-attention-2d-convolutional-neural-1","repo_url":"https://github.com/poifull10/PervasiveAttentionNet","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null},{"paper_slug":"pervasive-attention-2d-convolutional-neural-1","repo_url":"https://github.com/tdiggelm/nn-experiments","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok"}}],"tasks":[{"task_slug":"decoder","task_name":"Decoder"},{"task_slug":"machine-translation","task_name":"Machine Translation"},{"task_slug":"translation","task_name":"Translation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/machine-translation-on-iwslt2015-english","task":"Machine Translation","dataset":"IWSLT2015 English-German","model":"Pervasive Attention","rank_in_archive_order":4,"of":8,"metrics":{"BLEU score":"27.99"},"uses_additional_data":false},{"leaderboard":"/sota/machine-translation-on-iwslt2015-german","task":"Machine Translation","dataset":"IWSLT2015 German-English","model":"Pervasive Attention","rank_in_archive_order":2,"of":15,"metrics":{"BLEU score":"34.18"},"uses_additional_data":false}],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=1808.03867","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}