{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/deep-attentional-structured-representation","title":"Deep Attentional Structured Representation Learning for Visual Recognition","arxiv_id":"1805.05389","date":"2018-05-14","proceeding":null,"authors":["Krishna Kanth Nakka","Mathieu Salzmann"],"abstract":"Structured representations, such as Bags of Words, VLAD and Fisher Vectors,\nhave proven highly effective to tackle complex visual recognition tasks. As\nsuch, they have recently been incorporated into deep architectures. However,\nwhile effective, the resulting deep structured representation learning\nstrategies typically aggregate local features from the entire image, ignoring\nthe fact that, in complex recognition tasks, some regions provide much more\ndiscriminative information than others.\n  In this paper, we introduce an attentional structured representation learning\nframework that incorporates an image-specific attention mechanism within the\naggregation process. Our framework learns to predict jointly the image class\nlabel and an attention map in an end-to-end fashion and without any other\nsupervision than the target label. As evidenced by our experiments, this\nconsistently outperforms attention-less structured representation learning and\nyields state-of-the-art results on standard scene recognition and fine-grained\ncategorization benchmarks.","url_abs":"http://arxiv.org/abs/1805.05389v1","url_pdf":"http://arxiv.org/pdf/1805.05389v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"deep-attentional-structured-representation","repo_url":"https://github.com/ahmedest61/vlad-buff","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"GPL-3.0"}}],"tasks":[{"task_slug":"representation-learning","task_name":"Representation Learning"},{"task_slug":"scene-recognition","task_name":"Scene Recognition"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}