{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/visual-transformers-token-based-image","title":"Visual Transformers: Token-based Image Representation and Processing for Computer Vision","arxiv_id":"2006.03677","date":"2020-06-05","proceeding":null,"authors":["Bichen Wu","Chenfeng Xu","Xiaoliang Dai","Alvin Wan","Peizhao Zhang","Zhicheng Yan","Masayoshi Tomizuka","Joseph Gonzalez","Kurt Keutzer","Peter Vajda"],"abstract":"Computer vision has achieved remarkable success by (a) representing images as uniformly-arranged pixel arrays and (b) convolving highly-localized features. However, convolutions treat all image pixels equally regardless of importance; explicitly model all concepts across all images, regardless of content; and struggle to relate spatially-distant concepts. In this work, we challenge this paradigm by (a) representing images as semantic visual tokens and (b) running transformers to densely model token relationships. Critically, our Visual Transformer operates in a semantic token space, judiciously attending to different image parts based on context. This is in sharp contrast to pixel-space transformers that require orders-of-magnitude more compute. Using an advanced training recipe, our VTs significantly outperform their convolutional counterparts, raising ResNet accuracy on ImageNet top-1 by 4.6 to 7 points while using fewer FLOPs and parameters. For semantic segmentation on LIP and COCO-stuff, VT-based feature pyramid networks (FPN) achieve 0.35 points higher mIoU while reducing the FPN module's FLOPs by 6.5x.","url_abs":"https://arxiv.org/abs/2006.03677v4","url_pdf":"https://arxiv.org/pdf/2006.03677v4.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"visual-transformers-token-based-image","repo_url":"https://github.com/Diksha942/Visual-Transformers","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok"}},{"paper_slug":"visual-transformers-token-based-image","repo_url":"https://github.com/ShivamRajSharma/Transformer-Architectures-From-Scratch","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"visual-transformers-token-based-image","repo_url":"https://github.com/aws-samples/amazon-sagemaker-visual-transformer","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT-0"}},{"paper_slug":"visual-transformers-token-based-image","repo_url":"https://github.com/chenfengxu714/YOGO","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"BSD-2-Clause"}},{"paper_slug":"visual-transformers-token-based-image","repo_url":"https://github.com/tahmid0007/VisualTransformers","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}},{"paper_slug":"visual-transformers-token-based-image","repo_url":"https://github.com/thijswerrij/Transformers-CBIS-DDSM","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}},{"paper_slug":"visual-transformers-token-based-image","repo_url":"https://github.com/NazirNayal8/visual-transformer","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"pytorch","reach":{"status":"ok"}},{"paper_slug":"visual-transformers-token-based-image","repo_url":"https://github.com/Oguzhanercan/Vision-Transformers","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"none","reach":{"status":"ok"}}],"tasks":[{"task_slug":"classification","task_name":"General Classification"},{"task_slug":"image-classification","task_name":"Image Classification"},{"task_slug":"semantic-segmentation","task_name":"Semantic Segmentation"}],"methods":[{"method_slug":"1x1-convolution","method_name":"1x1 Convolution"},{"method_slug":"absolute-position-encodings","method_name":"Absolute Position Encodings"},{"method_slug":"adam","method_name":"Adam"},{"method_slug":"attention","method_name":"Attention"},{"method_slug":"average-pooling","method_name":"Average Pooling"},{"method_slug":"bpe","method_name":"BPE"},{"method_slug":"batch-normalization","method_name":"Batch Normalization"},{"method_slug":"bottleneck-residual-block","method_name":"Bottleneck Residual Block"},{"method_slug":"convolution","method_name":"Convolution"},{"method_slug":"dense-connections","method_name":"Dense Connections"},{"method_slug":"dropout","method_name":"Dropout"},{"method_slug":"fpn","method_name":"FPN"},{"method_slug":"global-average-pooling","method_name":"Global Average Pooling"},{"method_slug":"kaiming-initialization","method_name":"Kaiming Initialization"},{"method_slug":"label-smoothing","method_name":"Label Smoothing"},{"method_slug":"layer-normalization","method_name":"Layer Normalization"},{"method_slug":"linear-layer","method_name":"Linear Layer"},{"method_slug":"max-pooling","method_name":"Max Pooling"},{"method_slug":"multi-head-attention","method_name":"Multi-Head Attention"},{"method_slug":"position-wise-feed-forward-layer","method_name":"Position-Wise Feed-Forward Layer"},{"method_slug":"relu","method_name":"ReLU"},{"method_slug":"residual-block","method_name":"Residual Block"},{"method_slug":"residual-connection","method_name":"Residual Connection"},{"method_slug":"softmax","method_name":"Softmax"},{"method_slug":"transformer","method_name":"Transformer"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2006.03677","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2006.03677"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/aws-samples/amazon-sagemaker-visual-transformer","reach":{"status":"ok","spdx":"MIT-0"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/NazirNayal8/visual-transformer","reach":{"status":"ok"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/thijswerrij/Transformers-CBIS-DDSM","reach":{"status":"ok"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/tahmid0007/VisualTransformers","reach":{"status":"ok"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/Diksha942/Visual-Transformers","reach":{"status":"ok"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/ShivamRajSharma/Transformer-Architectures-From-Scratch","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/chenfengxu714/YOGO","reach":{"status":"ok","spdx":"BSD-2-Clause"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/Oguzhanercan/Vision-Transformers","reach":{"status":"ok"}}],"summary":{"ran":2,"unverified":4},"by_repo_kind":{"listed":{"samples":6,"ran":2,"repositories":2}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"c360bb211fd2a25c","entry":"conv1x1_1d","repo":"chenfengxu714/YOGO","repo_kind":"listed","path":"models/s3dis/yogo.py","file_url":"https://github.com/chenfengxu714/YOGO/blob/HEAD/models/s3dis/yogo.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"BSD-2-Clause","inline_ok":true,"mcp_get_code":{"code_sha256":"c360bb211fd2a25c"}},{"code_sha256_prefix":"3ae68d06c43d9791","entry":"get_box_corners_3d","repo":"chenfengxu714/YOGO","repo_kind":"listed","path":"modules/frustum.py","file_url":"https://github.com/chenfengxu714/YOGO/blob/HEAD/modules/frustum.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"BSD-2-Clause","inline_ok":true,"mcp_get_code":{"code_sha256":"3ae68d06c43d9791"}},{"code_sha256_prefix":"105fad12a622848b","entry":"augmentation_transform","repo":"chenfengxu714/YOGO","repo_kind":"listed","path":"datasets/shapenet.py","file_url":"https://github.com/chenfengxu714/YOGO/blob/HEAD/datasets/shapenet.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"BSD-2-Clause","inline_ok":true,"mcp_get_code":{"code_sha256":"105fad12a622848b"}},{"code_sha256_prefix":"50898529741237cc","entry":"conv1x1","repo":"chenfengxu714/YOGO","repo_kind":"listed","path":"models/s3dis/yogo.py","file_url":"https://github.com/chenfengxu714/YOGO/blob/HEAD/models/s3dis/yogo.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"BSD-2-Clause","inline_ok":true,"mcp_get_code":{"code_sha256":"50898529741237cc"}},{"code_sha256_prefix":"4a3aa44eded8d09b","entry":"eval_model","repo":"aws-samples/amazon-sagemaker-visual-transformer","repo_kind":"listed","path":"image-classification/code/vt-resnet-34.py","file_url":"https://github.com/aws-samples/amazon-sagemaker-visual-transformer/blob/HEAD/image-classification/code/vt-resnet-34.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT-0","inline_ok":true,"mcp_get_code":{"code_sha256":"4a3aa44eded8d09b"}},{"code_sha256_prefix":"dd70f0566a32d565","entry":"train_epoch","repo":"aws-samples/amazon-sagemaker-visual-transformer","repo_kind":"listed","path":"image-classification/code/vt-resnet-34.py","file_url":"https://github.com/aws-samples/amazon-sagemaker-visual-transformer/blob/HEAD/image-classification/code/vt-resnet-34.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT-0","inline_ok":true,"mcp_get_code":{"code_sha256":"dd70f0566a32d565"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}