{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/global-and-local-attention-networks-for","title":"Learning what and where to attend","arxiv_id":"1805.08819","date":"2018-05-22","proceeding":null,"authors":["Drew Linsley","Dan Shiebler","Sven Eberhardt","Thomas Serre"],"abstract":"Most recent gains in visual recognition have originated from the inclusion of attention mechanisms in deep convolutional networks (DCNs). Because these networks are optimized for object recognition, they learn where to attend using only a weak form of supervision derived from image class labels. Here, we demonstrate the benefit of using stronger supervisory signals by teaching DCNs to attend to image regions that humans deem important for object recognition. We first describe a large-scale online experiment (ClickMe) used to supplement ImageNet with nearly half a million human-derived \"top-down\" attention maps. Using human psychophysics, we confirm that the identified top-down features from ClickMe are more diagnostic than \"bottom-up\" saliency features for rapid image categorization. As a proof of concept, we extend a state-of-the-art attention network and demonstrate that adding ClickMe supervision significantly improves its accuracy and yields visual features that are more interpretable and more similar to those used by human observers.","url_abs":"https://arxiv.org/abs/1805.08819v4","url_pdf":"https://arxiv.org/pdf/1805.08819v4.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"global-and-local-attention-networks-for","repo_url":"https://github.com/serre-lab/harmonization","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"diagnostic","task_name":"Diagnostic"},{"task_slug":"image-categorization","task_name":"Image Categorization"},{"task_slug":"object-recognition","task_name":"Object Recognition"}],"methods":[{"method_slug":"gala","method_name":"GALA"}],"datasets_introduced":[],"methods_introduced":[{"slug":"gala","name":"GALA","full_name":"Global-and-Local attention"}],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1805.08819","atlas_url":"https://app.syntology.ai/?focus=1805.08819","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}