{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/context-aware-attentional-pooling-cap-for","title":"Context-aware Attentional Pooling (CAP) for Fine-grained Visual Classification","arxiv_id":"2101.06635","date":"2021-01-17","proceeding":null,"authors":["Ardhendu Behera","Zachary Wharton","Pradeep Hewage","Asish Bera"],"abstract":"Deep convolutional neural networks (CNNs) have shown a strong ability in mining discriminative object pose and parts information for image recognition. For fine-grained recognition, context-aware rich feature representation of object/scene plays a key role since it exhibits a significant variance in the same subcategory and subtle variance among different subcategories. Finding the subtle variance that fully characterizes the object/scene is not straightforward. To address this, we propose a novel context-aware attentional pooling (CAP) that effectively captures subtle changes via sub-pixel gradients, and learns to attend informative integral regions and their importance in discriminating different subcategories without requiring the bounding-box and/or distinguishable part annotations. We also introduce a novel feature encoding by considering the intrinsic consistency between the informativeness of the integral regions and their spatial structures to capture the semantic correlation among them. Our approach is simple yet extremely effective and can be easily applied on top of a standard classification backbone network. We evaluate our approach using six state-of-the-art (SotA) backbone networks and eight benchmark datasets. Our method significantly outperforms the SotA approaches on six datasets and is very competitive with the remaining two.","url_abs":"https://arxiv.org/abs/2101.06635v1","url_pdf":"https://arxiv.org/pdf/2101.06635v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"context-aware-attentional-pooling-cap-for","repo_url":"https://github.com/ArdhenduBehera/cap","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":null}],"tasks":[{"task_slug":"fine-grained-image-classification","task_name":"Fine-Grained Image Classification"},{"task_slug":"classification","task_name":"General Classification"},{"task_slug":"informativeness","task_name":"Informativeness"},{"task_slug":"object","task_name":"Object"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/fine-grained-image-classification-on-cub-200-1","task":"Fine-Grained Image Classification","dataset":"CUB-200-2011","model":"CAP","rank_in_archive_order":4,"of":30,"metrics":{"Accuracy":"91.8"},"uses_additional_data":false},{"leaderboard":"/sota/fine-grained-image-classification-on-fgvc","task":"Fine-Grained Image Classification","dataset":"FGVC Aircraft","model":"CAP","rank_in_archive_order":4,"of":57,"metrics":{"Accuracy":"94.9%","PARAMS":"34.2"},"uses_additional_data":false},{"leaderboard":"/sota/fine-grained-image-classification-on-food-101","task":"Fine-Grained Image Classification","dataset":"Food-101","model":"CAP","rank_in_archive_order":1,"of":15,"metrics":{"Accuracy":"98.6","PARAMS":"34.2"},"uses_additional_data":false},{"leaderboard":"/sota/fine-grained-image-classification-on-nabirds","task":"Fine-Grained Image Classification","dataset":"NABirds","model":"CAP","rank_in_archive_order":13,"of":30,"metrics":{"Accuracy":"91.0%"},"uses_additional_data":false},{"leaderboard":"/sota/fine-grained-image-classification-on-stanford","task":"Fine-Grained Image Classification","dataset":"Stanford Cars","model":"CAP","rank_in_archive_order":10,"of":83,"metrics":{"Accuracy":"95.7%"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2101.06635","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2101.06635"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/ArdhenduBehera/cap","reach":null}],"summary":{"unverified":2},"by_repo_kind":{"listed":{"samples":2,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"6fa32f316d84dd85","entry":"SelfAttention","repo":"ArdhenduBehera/cap","repo_kind":"listed","path":"SelfAttention.py","file_url":"https://github.com/ArdhenduBehera/cap/blob/HEAD/SelfAttention.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"6fa32f316d84dd85"}},{"code_sha256_prefix":"e706320ae0b6b7a2","entry":"hw_flatten","repo":"ArdhenduBehera/cap","repo_kind":"listed","path":"SelfAttention.py","file_url":"https://github.com/ArdhenduBehera/cap/blob/HEAD/SelfAttention.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"e706320ae0b6b7a2"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}