{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/sca-cnn-spatial-and-channel-wise-attention-in","title":"SCA-CNN: Spatial and Channel-wise Attention in Convolutional Networks for Image Captioning","arxiv_id":"1611.05594","date":"2016-11-17","proceeding":"CVPR 2017 7","authors":["Long Chen","Hanwang Zhang","Jun Xiao","Liqiang Nie","Jian Shao","Wei Liu","Tat-Seng Chua"],"abstract":"Visual attention has been successfully applied in structural prediction tasks\nsuch as visual captioning and question answering. Existing visual attention\nmodels are generally spatial, i.e., the attention is modeled as spatial\nprobabilities that re-weight the last conv-layer feature map of a CNN encoding\nan input image. However, we argue that such spatial attention does not\nnecessarily conform to the attention mechanism --- a dynamic feature extractor\nthat combines contextual fixations over time, as CNN features are naturally\nspatial, channel-wise and multi-layer. In this paper, we introduce a novel\nconvolutional neural network dubbed SCA-CNN that incorporates Spatial and\nChannel-wise Attentions in a CNN. In the task of image captioning, SCA-CNN\ndynamically modulates the sentence generation context in multi-layer feature\nmaps, encoding where (i.e., attentive spatial locations at multiple layers) and\nwhat (i.e., attentive channels) the visual attention is. We evaluate the\nproposed SCA-CNN architecture on three benchmark image captioning datasets:\nFlickr8K, Flickr30K, and MSCOCO. It is consistently observed that SCA-CNN\nsignificantly outperforms state-of-the-art visual attention-based image\ncaptioning methods.","url_abs":"http://arxiv.org/abs/1611.05594v2","url_pdf":"http://arxiv.org/pdf/1611.05594v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"sca-cnn-spatial-and-channel-wise-attention-in","repo_url":"https://github.com/zjuchenlong/sca-cnn","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok"}},{"paper_slug":"sca-cnn-spatial-and-channel-wise-attention-in","repo_url":"https://github.com/HaleyPei/Implement-of-SCA-CNN","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"image-captioning","task_name":"Image Captioning"},{"task_slug":"sentence","task_name":"Sentence"}],"methods":[{"method_slug":"sca-cnn","method_name":"SCA-CNN"}],"datasets_introduced":[],"methods_introduced":[{"slug":"sca-cnn","name":"SCA-CNN","full_name":"Spatial and Channel-wise Attention-based Convolutional Neural Network"}],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1611.05594","atlas_url":"https://app.syntology.ai/?focus=1611.05594","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}