{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/glit-neural-architecture-search-for-global","title":"GLiT: Neural Architecture Search for Global and Local Image Transformer","arxiv_id":"2107.02960","date":"2021-07-07","proceeding":"ICCV 2021 10","authors":["BoYu Chen","Peixia Li","Chuming Li","Baopu Li","Lei Bai","Chen Lin","Ming Sun","Junjie Yan","Wanli Ouyang"],"abstract":"We introduce the first Neural Architecture Search (NAS) method to find a better transformer architecture for image recognition. Recently, transformers without CNN-based backbones are found to achieve impressive performance for image recognition. However, the transformer is designed for NLP tasks and thus could be sub-optimal when directly used for image recognition. In order to improve the visual representation ability for transformers, we propose a new search space and searching algorithm. Specifically, we introduce a locality module that models the local correlations in images explicitly with fewer computational cost. With the locality module, our search space is defined to let the search algorithm freely trade off between global and local information as well as optimizing the low-level design choice in each module. To tackle the problem caused by huge search space, a hierarchical neural architecture search method is proposed to search the optimal vision transformer from two levels separately with the evolutionary algorithm. Extensive experiments on the ImageNet dataset demonstrate that our method can find more discriminative and efficient transformer variants than the ResNet family (e.g., ResNet101) and the baseline ViT for image classification.","url_abs":"https://arxiv.org/abs/2107.02960v3","url_pdf":"https://arxiv.org/pdf/2107.02960v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"glit-neural-architecture-search-for-global","repo_url":"https://github.com/bychen515/glit","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}},{"paper_slug":"glit-neural-architecture-search-for-global","repo_url":"https://github.com/lpxtt/simtrack","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"image-classification","task_name":"Image Classification"},{"task_slug":"architecture-search","task_name":"Neural Architecture Search"},{"task_slug":"image-classification","task_name":"image-classification"}],"methods":[{"method_slug":"1x1-convolution","method_name":"1x1 Convolution"},{"method_slug":"attention","method_name":"Attention"},{"method_slug":"average-pooling","method_name":"Average Pooling"},{"method_slug":"batch-normalization","method_name":"Batch Normalization"},{"method_slug":"bottleneck-residual-block","method_name":"Bottleneck Residual Block"},{"method_slug":"convolution","method_name":"Convolution"},{"method_slug":"dense-connections","method_name":"Dense Connections"},{"method_slug":"global-average-pooling","method_name":"Global Average Pooling"},{"method_slug":"kaiming-initialization","method_name":"Kaiming Initialization"},{"method_slug":"layer-normalization","method_name":"Layer Normalization"},{"method_slug":"linear-layer","method_name":"Linear Layer"},{"method_slug":"max-pooling","method_name":"Max Pooling"},{"method_slug":"multi-head-attention","method_name":"Multi-Head Attention"},{"method_slug":"relu","method_name":"ReLU"},{"method_slug":"residual-block","method_name":"Residual Block"},{"method_slug":"residual-connection","method_name":"Residual Connection"},{"method_slug":"softmax","method_name":"Softmax"},{"method_slug":"vision-transformer","method_name":"Vision Transformer"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/image-classification-on-imagenet","task":"Image Classification","dataset":"ImageNet","model":"GLiT-Bases","rank_in_archive_order":555,"of":1060,"metrics":{"GFLOPs":"17","Number of params":"96.1M","Top 1 Accuracy":"82.3%"},"uses_additional_data":false},{"leaderboard":"/sota/image-classification-on-imagenet","task":"Image Classification","dataset":"ImageNet","model":"GLiT-Smalls","rank_in_archive_order":699,"of":1060,"metrics":{"GFLOPs":"4.4","Number of params":"24.6M","Top 1 Accuracy":"80.5%"},"uses_additional_data":false},{"leaderboard":"/sota/image-classification-on-imagenet","task":"Image Classification","dataset":"ImageNet","model":"GLiT-Tinys","rank_in_archive_order":920,"of":1060,"metrics":{"GFLOPs":"1.4","Number of params":"7.2M","Top 1 Accuracy":"76.3%"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2107.02960","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2107.02960"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/bychen515/glit","reach":{"status":"ok"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/lpxtt/simtrack","reach":null},{"provenance":"deterministic:regex_extraction","url":"https://github.com/bychen515/GLiT","reach":{"status":"ok"}}],"summary":{"ran":2,"unverified":2},"by_repo_kind":{"listed":{"samples":4,"ran":2,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"185ba7199442da8a","entry":"TransformerEncoder","repo":"lpxtt/simtrack","repo_kind":"listed","path":"lib/models/stark/transformer.py","file_url":"https://github.com/lpxtt/simtrack/blob/HEAD/lib/models/stark/transformer.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"185ba7199442da8a"}},{"code_sha256_prefix":"0fa568810a644276","entry":"TransformerEncoderLayer","repo":"lpxtt/simtrack","repo_kind":"listed","path":"lib/models/stark/transformer.py","file_url":"https://github.com/lpxtt/simtrack/blob/HEAD/lib/models/stark/transformer.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"0fa568810a644276"}},{"code_sha256_prefix":"40172c198ba18de8","entry":"Transformer","repo":"lpxtt/simtrack","repo_kind":"listed","path":"lib/models/stark/transformer.py","file_url":"https://github.com/lpxtt/simtrack/blob/HEAD/lib/models/stark/transformer.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"40172c198ba18de8"}},{"code_sha256_prefix":"45c535d5c49fc225","entry":"TransformerDecoderLayer","repo":"lpxtt/simtrack","repo_kind":"listed","path":"lib/models/stark/transformer.py","file_url":"https://github.com/lpxtt/simtrack/blob/HEAD/lib/models/stark/transformer.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"45c535d5c49fc225"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}