{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/demystify-transformers-convolutions-in-modern","title":"Demystify Transformers & Convolutions in Modern Image Deep Networks","arxiv_id":"2211.05781","date":"2022-11-10","proceeding":null,"authors":["Xiaowei Hu","Min Shi","Weiyun Wang","Sitong Wu","Linjie Xing","Wenhai Wang","Xizhou Zhu","Lewei Lu","Jie zhou","Xiaogang Wang","Yu Qiao","Jifeng Dai"],"abstract":"Vision transformers have gained popularity recently, leading to the development of new vision backbones with improved features and consistent performance gains. However, these advancements are not solely attributable to novel feature transformation designs; certain benefits also arise from advanced network-level and block-level architectures. This paper aims to identify the real gains of popular convolution and attention operators through a detailed study. We find that the key difference among these feature transformation modules, such as attention or convolution, lies in their spatial feature aggregation approach, known as the \"spatial token mixer\" (STM). To facilitate an impartial comparison, we introduce a unified architecture to neutralize the impact of divergent network-level and block-level designs. Subsequently, various STMs are integrated into this unified framework for comprehensive comparative analysis. Our experiments on various tasks and an analysis of inductive bias show a significant performance boost due to advanced network-level and block-level designs, but performance differences persist among different STMs. Our detailed analysis also reveals various findings about different STMs, including effective receptive fields, invariance, and adversarial robustness tests.","url_abs":"https://arxiv.org/abs/2211.05781v3","url_pdf":"https://arxiv.org/pdf/2211.05781v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"demystify-transformers-convolutions-in-modern","repo_url":"https://github.com/opengvlab/stm-evaluation","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"adversarial-robustness","task_name":"Adversarial Robustness"},{"task_slug":"image-deep-networks","task_name":"Image Deep Networks"},{"task_slug":"spatial-token-mixer","task_name":"Spatial Token Mixer"}],"methods":[{"method_slug":"absolute-position-encodings","method_name":"Absolute Position Encodings"},{"method_slug":"adam","method_name":"Adam"},{"method_slug":"attention","method_name":"Attention"},{"method_slug":"bpe","method_name":"BPE"},{"method_slug":"convolution","method_name":"Convolution"},{"method_slug":"dense-connections","method_name":"Dense Connections"},{"method_slug":"dropout","method_name":"Dropout"},{"method_slug":"label-smoothing","method_name":"Label Smoothing"},{"method_slug":"layer-normalization","method_name":"Layer Normalization"},{"method_slug":"linear-layer","method_name":"Linear Layer"},{"method_slug":"multi-head-attention","method_name":"Multi-Head Attention"},{"method_slug":"position-wise-feed-forward-layer","method_name":"Position-Wise Feed-Forward Layer"},{"method_slug":"residual-connection","method_name":"Residual Connection"},{"method_slug":"softmax","method_name":"Softmax"},{"method_slug":"transformer","method_name":"Transformer"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2211.05781","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2211.05781"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/opengvlab/stm-evaluation","reach":{"status":"ok","spdx":"MIT"}}],"summary":{"ran":1,"ran_honours":1,"unverified":3},"by_repo_kind":{"official":{"samples":5,"ran":2,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"6a0317bfd6be9557","entry":"build_transform","repo":"opengvlab/stm-evaluation","repo_kind":"official","path":"classification/datasets.py","file_url":"https://github.com/opengvlab/stm-evaluation/blob/HEAD/classification/datasets.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"6a0317bfd6be9557"}},{"code_sha256_prefix":"24400bc088ed228d","entry":"make_divisible","repo":"opengvlab/stm-evaluation","repo_kind":"official","path":"classification/models/blocks/halonet.py","file_url":"https://github.com/opengvlab/stm-evaluation/blob/HEAD/classification/models/blocks/halonet.py","link_basis":"harvester_set","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":"well_formed","behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"24400bc088ed228d"}},{"code_sha256_prefix":"b700fb0a3496b255","entry":"cosine_scheduler","repo":"opengvlab/stm-evaluation","repo_kind":"official","path":"classification/utils.py","file_url":"https://github.com/opengvlab/stm-evaluation/blob/HEAD/classification/utils.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"b700fb0a3496b255"}},{"code_sha256_prefix":"199223de2f9704cf","entry":"get_num_layer_for_convnext","repo":"opengvlab/stm-evaluation","repo_kind":"official","path":"classification/optim_factory.py","file_url":"https://github.com/opengvlab/stm-evaluation/blob/HEAD/classification/optim_factory.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"199223de2f9704cf"}},{"code_sha256_prefix":"0ef0019514eaaa09","entry":"get_parameter_groups","repo":"opengvlab/stm-evaluation","repo_kind":"official","path":"classification/optim_factory.py","file_url":"https://github.com/opengvlab/stm-evaluation/blob/HEAD/classification/optim_factory.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"0ef0019514eaaa09"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}