{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/group-contextualization-for-video-recognition","title":"Group Contextualization for Video Recognition","arxiv_id":"2203.09694","date":"2022-03-18","proceeding":"CVPR 2022 1","authors":["Yanbin Hao","Hao Zhang","Chong-Wah Ngo","Xiangnan He"],"abstract":"Learning discriminative representation from the complex spatio-temporal dynamic space is essential for video recognition. On top of those stylized spatio-temporal computational units, further refining the learnt feature with axial contexts is demonstrated to be promising in achieving this goal. However, previous works generally focus on utilizing a single kind of contexts to calibrate entire feature channels and could hardly apply to deal with diverse video activities. The problem can be tackled by using pair-wise spatio-temporal attentions to recompute feature response with cross-axis contexts at the expense of heavy computations. In this paper, we propose an efficient feature refinement method that decomposes the feature channels into several groups and separately refines them with different axial contexts in parallel. We refer this lightweight feature calibration as group contextualization (GC). Specifically, we design a family of efficient element-wise calibrators, i.e., ECal-G/S/T/L, where their axial contexts are information dynamics aggregated from other axes either globally or locally, to contextualize feature channel groups. The GC module can be densely plugged into each residual layer of the off-the-shelf video networks. With little computational overhead, consistent improvement is observed when plugging in GC on different networks. By utilizing calibrators to embed feature with four different kinds of contexts in parallel, the learnt representation is expected to be more resilient to diverse types of activities. On videos with rich temporal variations, empirically GC can boost the performance of 2D-CNN (e.g., TSN and TSM) to a level comparable to the state-of-the-art video networks. Code is available at https://github.com/haoyanbin918/Group-Contextualization.","url_abs":"https://arxiv.org/abs/2203.09694v1","url_pdf":"https://arxiv.org/pdf/2203.09694v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"group-contextualization-for-video-recognition","repo_url":"https://github.com/haoyanbin918/group-contextualization","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"egocentric-activity-recognition","task_name":"Egocentric Activity Recognition"},{"task_slug":"video-recognition","task_name":"Video Recognition"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/action-recognition-on-diving-48","task":"Action Recognition","dataset":"Diving-48","model":"GC-TDN","rank_in_archive_order":8,"of":18,"metrics":{"Accuracy":"87.6"},"uses_additional_data":false},{"leaderboard":"/sota/action-recognition-in-videos-on-something","task":"Action Recognition","dataset":"Something-Something V2","model":"GC-TDN Ensemble (R50,8+16)","rank_in_archive_order":58,"of":123,"metrics":{"GFLOPs":"110.1","Parameters":"27.4","Top-1 Accuracy":"67.8","Top-5 Accuracy":"91.2"},"uses_additional_data":true},{"leaderboard":"/sota/egocentric-activity-recognition-on-egtea-1","task":"Egocentric Activity Recognition","dataset":"EGTEA","model":"GC-TSM","rank_in_archive_order":3,"of":6,"metrics":{"Average Accuracy":"65.1"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2203.09694","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2203.09694"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"deterministic:regex_extraction","url":"https://github.com/haoyanbin918/Group-Contextualization","reach":null}],"summary":{"ran":5},"by_repo_kind":{"official":{"samples":5,"ran":5,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"681847d2c90069fd","entry":"Bottleneck","repo":"haoyanbin918/Group-Contextualization","repo_kind":"official","path":"nets/GC_GST.py","file_url":"https://github.com/haoyanbin918/Group-Contextualization/blob/HEAD/nets/GC_GST.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"681847d2c90069fd"}},{"code_sha256_prefix":"bfd0fc32160812fb","entry":"GC_CLLDnb","repo":"haoyanbin918/Group-Contextualization","repo_kind":"official","path":"nets/GC_GST.py","file_url":"https://github.com/haoyanbin918/Group-Contextualization/blob/HEAD/nets/GC_GST.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"bfd0fc32160812fb"}},{"code_sha256_prefix":"ad22890d898cf1b4","entry":"GC_L33Dnb","repo":"haoyanbin918/Group-Contextualization","repo_kind":"official","path":"nets/GC_GST.py","file_url":"https://github.com/haoyanbin918/Group-Contextualization/blob/HEAD/nets/GC_GST.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"ad22890d898cf1b4"}},{"code_sha256_prefix":"b46bb46ae87e6dd5","entry":"GC_S23DDnb","repo":"haoyanbin918/Group-Contextualization","repo_kind":"official","path":"nets/GC_GST.py","file_url":"https://github.com/haoyanbin918/Group-Contextualization/blob/HEAD/nets/GC_GST.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"b46bb46ae87e6dd5"}},{"code_sha256_prefix":"b86f2ed4b012b0fa","entry":"GC_T13Dnb","repo":"haoyanbin918/Group-Contextualization","repo_kind":"official","path":"nets/GC_GST.py","file_url":"https://github.com/haoyanbin918/Group-Contextualization/blob/HEAD/nets/GC_GST.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"b86f2ed4b012b0fa"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}