{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/attentional-pooling-for-action-recognition","title":"Attentional Pooling for Action Recognition","arxiv_id":"1711.01467","date":"2017-11-04","proceeding":"NeurIPS 2017 12","authors":["Rohit Girdhar","Deva Ramanan"],"abstract":"We introduce a simple yet surprisingly powerful model to incorporate\nattention in action recognition and human object interaction tasks. Our\nproposed attention module can be trained with or without extra supervision, and\ngives a sizable boost in accuracy while keeping the network size and\ncomputational cost nearly the same. It leads to significant improvements over\nstate of the art base architecture on three standard action recognition\nbenchmarks across still images and videos, and establishes new state of the art\non MPII dataset with 12.5% relative improvement. We also perform an extensive\nanalysis of our attention module both empirically and analytically. In terms of\nthe latter, we introduce a novel derivation of bottom-up and top-down attention\nas low-rank approximations of bilinear pooling methods (typically used for\nfine-grained classification). From this perspective, our attention formulation\nsuggests a novel characterization of action recognition as a fine-grained\nrecognition problem.","url_abs":"http://arxiv.org/abs/1711.01467v3","url_pdf":"http://arxiv.org/pdf/1711.01467v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"attentional-pooling-for-action-recognition","repo_url":"https://github.com/rohitgirdhar/AttentionalPoolingAction","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok","spdx":"NOASSERTION"}}],"tasks":[{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"human-object-interaction-detection","task_name":"Human-Object Interaction Detection"},{"task_slug":"action-recognition","task_name":"Temporal Action Localization"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/human-object-interaction-detection-on-hico-1","task":"Human-Object Interaction Detection","dataset":"HICO","model":"Girdhar & Ramanan","rank_in_archive_order":7,"of":8,"metrics":{"mAP":"34.6"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/1711.01467","atlas_url":"https://app.syntology.ai/?focus=1711.01467","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}