{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/spectformer-frequency-and-attention-is-what","title":"SpectFormer: Frequency and Attention is what you need in a Vision Transformer","arxiv_id":"2304.06446","date":"2023-04-13","proceeding":null,"authors":["Badri N. Patro","Vinay P. Namboodiri","Vijay Srinivas Agneeswaran"],"abstract":"Vision transformers have been applied successfully for image recognition tasks. There have been either multi-headed self-attention based (ViT \\cite{dosovitskiy2020image}, DeIT, \\cite{touvron2021training}) similar to the original work in textual models or more recently based on spectral layers (Fnet\\cite{lee2021fnet}, GFNet\\cite{rao2021global}, AFNO\\cite{guibas2021efficient}). We hypothesize that both spectral and multi-headed attention plays a major role. We investigate this hypothesis through this work and observe that indeed combining spectral and multi-headed attention layers provides a better transformer architecture. We thus propose the novel Spectformer architecture for transformers that combines spectral and multi-headed attention layers. We believe that the resulting representation allows the transformer to capture the feature representation appropriately and it yields improved performance over other transformer representations. For instance, it improves the top-1 accuracy by 2\\% on ImageNet compared to both GFNet-H and LiT. SpectFormer-S reaches 84.25\\% top-1 accuracy on ImageNet-1K (state of the art for small version). Further, Spectformer-L achieves 85.7\\% that is the state of the art for the comparable base version of the transformers. We further ensure that we obtain reasonable results in other scenarios such as transfer learning on standard datasets such as CIFAR-10, CIFAR-100, Oxford-IIIT-flower, and Standford Car datasets. We then investigate its use in downstream tasks such of object detection and instance segmentation on the MS-COCO dataset and observe that Spectformer shows consistent performance that is comparable to the best backbones and can be further optimized and improved. Hence, we believe that combined spectral and attention layers are what are needed for vision transformers.","url_abs":"https://arxiv.org/abs/2304.06446v2","url_pdf":"https://arxiv.org/pdf/2304.06446v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"spectformer-frequency-and-attention-is-what","repo_url":"https://github.com/Hazqeel09/ellzaf_ml","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"instance-segmentation","task_name":"Instance Segmentation"},{"task_slug":"object-detection","task_name":"Object Detection"},{"task_slug":"semantic-segmentation","task_name":"Semantic Segmentation"},{"task_slug":"transfer-learning","task_name":"Transfer Learning"},{"task_slug":"object-detection-1","task_name":"object-detection"}],"methods":[{"method_slug":"base","method_name":"BASE"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2304.06446","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2304.06446"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/Hazqeel09/ellzaf_ml","reach":null}],"summary":{"ran_draft_wrong":3,"ran_violates":1,"ran_fixture":1},"by_repo_kind":{"listed":{"samples":4,"ran":4,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":1,"samples":[{"code_sha256_prefix":"4a4c555b7cc644cc","entry":"build_finetune_optimizer","repo":"Hazqeel09/ellzaf_ml","repo_kind":"listed","path":"ellzaf_ml/tools/build_optimizer.py","file_url":"https://github.com/Hazqeel09/ellzaf_ml/blob/HEAD/ellzaf_ml/tools/build_optimizer.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"4a4c555b7cc644cc"}},{"code_sha256_prefix":"fb8ed2605c193c08","entry":"build_pretrain_optimizer","repo":"Hazqeel09/ellzaf_ml","repo_kind":"listed","path":"ellzaf_ml/tools/build_optimizer.py","file_url":"https://github.com/Hazqeel09/ellzaf_ml/blob/HEAD/ellzaf_ml/tools/build_optimizer.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"fb8ed2605c193c08"}},{"code_sha256_prefix":"2089fe788b366d85","entry":"get_pretrain_param_groups","repo":"Hazqeel09/ellzaf_ml","repo_kind":"listed","path":"ellzaf_ml/tools/build_optimizer.py","file_url":"https://github.com/Hazqeel09/ellzaf_ml/blob/HEAD/ellzaf_ml/tools/build_optimizer.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"2089fe788b366d85"}},{"code_sha256_prefix":"6ba8cee9f5daea41","entry":"pair","repo":null,"repo_kind":null,"path":null,"file_url":null,"link_basis":"identical_code_first_harvested_elsewhere","language":"python","status":"ran_violates","verification_level":1,"contract_check":"VIOLATES","metamorphic_tier":"well_formed","behaviour_fingerprint":true,"licence":null,"inline_ok":false,"mcp_get_code":{"code_sha256":"6ba8cee9f5daea41"}},{"code_sha256_prefix":"17f6d223949a7e58","entry":"posemb_sincos_2d","repo":"Hazqeel09/ellzaf_ml","repo_kind":"listed","path":"ellzaf_ml/models/spectformer.py","file_url":"https://github.com/Hazqeel09/ellzaf_ml/blob/HEAD/ellzaf_ml/models/spectformer.py","link_basis":"first_harvest_node","language":"python","status":"ran_fixture","verification_level":1,"contract_check":"RAISES","metamorphic_tier":"invariant","behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"17f6d223949a7e58"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}