{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/uninext-exploring-a-unified-architecture-for","title":"UniNeXt: Exploring A Unified Architecture for Vision Recognition","arxiv_id":"2304.13700","date":"2023-04-26","proceeding":null,"authors":["Fangjian Lin","Jianlong Yuan","Sitong Wu","Fan Wang","Zhibin Wang"],"abstract":"Vision Transformers have shown great potential in computer vision tasks. Most recent works have focused on elaborating the spatial token mixer for performance gains. However, we observe that a well-designed general architecture can significantly improve the performance of the entire backbone, regardless of which spatial token mixer is equipped. In this paper, we propose UniNeXt, an improved general architecture for the vision backbone. To verify its effectiveness, we instantiate the spatial token mixer with various typical and modern designs, including both convolution and attention modules. Compared with the architecture in which they are first proposed, our UniNeXt architecture can steadily boost the performance of all the spatial token mixers, and narrows the performance gap among them. Surprisingly, our UniNeXt equipped with naive local window attention even outperforms the previous state-of-the-art. Interestingly, the ranking of these spatial token mixers also changes under our UniNeXt, suggesting that an excellent spatial token mixer may be stifled due to a suboptimal general architecture, which further shows the importance of the study on the general architecture of vision backbone.","url_abs":"https://arxiv.org/abs/2304.13700v3","url_pdf":"https://arxiv.org/pdf/2304.13700v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"uninext-exploring-a-unified-architecture-for","repo_url":"https://github.com/jianlong-yuan/uninext","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"spatial-token-mixer","task_name":"Spatial Token Mixer"}],"methods":[{"method_slug":"convolution","method_name":"Convolution"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/2304.13700","atlas_url":"https://app.syntology.ai/?focus=2304.13700","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2304.13700"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-25T09:33:49+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/jianlong-yuan/uninext","reach":null}],"summary":{"ran_fixture":3},"by_repo_kind":{"official":{"samples":3,"ran":3,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":3,"samples":[{"code_sha256_prefix":"e83a7e641ac2b922","entry":"group2image","repo":"jianlong-yuan/uninext","repo_kind":"official","path":"Classification/models/UniNeXt.py","file_url":"https://github.com/jianlong-yuan/uninext/blob/HEAD/Classification/models/UniNeXt.py","link_basis":"first_harvest_node","language":"python","status":"ran_fixture","verification_level":1,"contract_check":"RAISES","metamorphic_tier":"invariant","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"e83a7e641ac2b922"}},{"code_sha256_prefix":"29dae49a66c72e53","entry":"img2group","repo":"jianlong-yuan/uninext","repo_kind":"official","path":"Classification/models/UniNeXt.py","file_url":"https://github.com/jianlong-yuan/uninext/blob/HEAD/Classification/models/UniNeXt.py","link_basis":"first_harvest_node","language":"python","status":"ran_fixture","verification_level":1,"contract_check":"RAISES","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"29dae49a66c72e53"}},{"code_sha256_prefix":"3177aba0ae8bd146","entry":"local_group","repo":"jianlong-yuan/uninext","repo_kind":"official","path":"Classification/models/UniNeXt.py","file_url":"https://github.com/jianlong-yuan/uninext/blob/HEAD/Classification/models/UniNeXt.py","link_basis":"first_harvest_node","language":"python","status":"ran_fixture","verification_level":1,"contract_check":"RAISES","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"3177aba0ae8bd146"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}