Methods › Computer Vision › Vision Transformers › Vision Transformer › Papers, page 22
Vision Transformer
Papers archive 2025-07-28
archive papers tagged: 2,144 · with a code link: 1,051 · where Syntology ran a sample: 328 (286 with a run with no instrument failure, 42 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (328 of 2,144 tagged: 286 with a run with no instrument failure, 42 where every run was a failure of Syntology's instrument)
Page 22 of 22: papers 2,101 to 2,144 of 2,144, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Segmenter: Transformer for Semantic Segmentation 12 May 2021 · 8 repositories · arXiv:2105.05633Syntology community repositories only · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 7 harvested samples)
-
Self-Supervised Learning with Swin Transformers 10 May 2021 · 6 repositories · arXiv:2105.04553
-
Instances as Queries 5 May 2021 · 5 repositories · arXiv:2105.01928Syntology community repositories only · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples) · 1 pointer-only (licence)
-
TransHash: Transformer-based Hamming Hashing for Efficient Image Retrieval 5 May 2021 · 0 repositories · arXiv:2105.01823
-
Emerging Properties in Self-Supervised Vision Transformers 29 Apr 2021 · 32 repositories · arXiv:2104.14294Syntology official: no sample here; runs from other or unrecorded repositories · 14 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 1 violated, 9 with no contract checked; 4 where Syntology's instrument failed) · 6 unverified (of 20 harvested samples) · 4 pointer-only (licence)
-
Twins: Revisiting the Design of Spatial Attention in Vision Transformers 28 Apr 2021 · 9 repositories · arXiv:2104.13840Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Vision Transformers with Patch Diversification 26 Apr 2021 · 1 repository · arXiv:2104.12753
-
Visual Saliency Transformer 25 Apr 2021 · 2 repositories · arXiv:2104.12099Syntology 7 ran (of which 4 constructed an object rather than computing a result; 7 with no instrument failure: 1 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 11 harvested samples) · 11 pointer-only (licence)
-
So-ViT: Mind Visual Tokens for Vision Transformer 22 Apr 2021 · 1 repository · arXiv:2104.10935
-
All Tokens Matter: Token Labeling for Training Better Vision Transformers 22 Apr 2021 · 7 repositories · arXiv:2104.10858Syntology official (archive's flag): 1 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 2 unverified (of 5 harvested samples) · 3 pointer-only (licence)
-
VATT: Transformers for Multimodal Self-Supervised Learning from Raw Video, Audio and Text 22 Apr 2021 · 5 repositories · arXiv:2104.11178Syntology community repositories only · 5 ran (of which 4 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 8 harvested samples) · 8 pointer-only (licence)
-
Generative Transformer for Accurate and Reliable Salient Object Detection 20 Apr 2021 · 2 repositories · arXiv:2104.10127
-
Vision Transformer Pruning 17 Apr 2021 · 2 repositories · arXiv:2104.08500Syntology 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples) · 1 pointer-only (licence)
-
Shoulder Implant X-Ray Manufacturer Classification: Exploring with Vision Transformer 15 Apr 2021 · 1 repository · arXiv:2104.07667
-
Vision Transformer using Low-level Chest X-ray Feature Corpus for COVID-19 Diagnosis and Severity Quantification 15 Apr 2021 · 0 repositories · arXiv:2104.07235
-
VTGAN: Semi-supervised Retinal Image Synthesis and Disease Prediction using Vision Transformers 14 Apr 2021 · 2 repositories · arXiv:2104.06757
-
ViT-V-Net: Vision Transformer for Unsupervised Volumetric Medical Image Registration 13 Apr 2021 · 1 repository · arXiv:2104.06468Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 2 honoured, 0 violated, 4 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 8 harvested samples)
-
Deepfake Detection Scheme Based on Vision Transformer and Distillation 3 Apr 2021 · 1 repository · arXiv:2104.01353
-
Putting NeRF on a Diet: Semantically Consistent Few-Shot View Synthesis 1 Apr 2021 · 2 repositories · arXiv:2104.00677Syntology official (archive's flag): 1 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 4 harvested samples) · 2 pointer-only (licence)
-
Rethinking Spatial Dimensions of Vision Transformers 30 Mar 2021 · 12 repositories · arXiv:2103.16302Syntology official: harvested, nothing ran · 10 ran (of which 6 constructed an object rather than computing a result; 10 with no instrument failure: 2 honoured, 0 violated, 8 with no contract checked; 0 where Syntology's instrument failed) · 10 unverified (of 20 harvested samples)
-
Multi-Scale Vision Longformer: A New Vision Transformer for High-Resolution Image Encoding 29 Mar 2021 · 3 repositories · arXiv:2103.15358Syntology official (archive's flag): 3 ran · 8 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 1 honoured, 2 violated, 2 with no contract checked; 3 where Syntology's instrument failed) · 2 unverified (of 10 harvested samples) · 2 pointer-only (licence)
-
CvT: Introducing Convolutions to Vision Transformers 29 Mar 2021 · 16 repositories · arXiv:2103.15808Syntology official (archive's flag): 10 ran · 39 ran (of which 19 constructed an object rather than computing a result; 36 with no instrument failure: 2 honoured, 0 violated, 34 with no contract checked; 3 where Syntology's instrument failed) · 8 unverified (of 47 harvested samples) · 8 pointer-only (licence)
-
CrossViT: Cross-Attention Multi-Scale Vision Transformer for Image Classification 27 Mar 2021 · 15 repositories · arXiv:2103.14899Syntology official (archive's flag): 4 ran · 17 ran (of which 10 constructed an object rather than computing a result; 17 with no instrument failure: 0 honoured, 1 violated, 16 with no contract checked; 0 where Syntology's instrument failed) · 9 unverified (of 26 harvested samples)
-
Understanding Robustness of Transformers for Image Classification 26 Mar 2021 · 0 repositories · arXiv:2103.14586
-
AutoMix: Unveiling the Power of Mixup for Stronger Classifiers 24 Mar 2021 · 3 repositories · arXiv:2103.13027Syntology community repositories only · 4 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 2 where Syntology's instrument failed) · 2 unverified (of 6 harvested samples) · 3 pointer-only (licence)
-
Vision Transformers for Dense Prediction 24 Mar 2021 · 15 repositories · arXiv:2103.13413Syntology 65 ran (of which 19 constructed an object rather than computing a result; 32 with no instrument failure: 0 honoured, 0 violated, 32 with no contract checked; 33 where Syntology's instrument failed) · 51 unverified (of 116 harvested samples) · 15 pointer-only (licence)
-
Danish Fungi 2020 -- Not Just Another Image Recognition Dataset 18 Mar 2021 · 1 repository · arXiv:2103.10107
-
TransFG: A Transformer Architecture for Fine-grained Recognition 14 Mar 2021 · 2 repositories · arXiv:2103.07976Syntology official (archive's flag): 1 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 2 pointer-only (licence)
-
Severity Quantification and Lesion Localization of COVID-19 on CXR using Vision Transformer 12 Mar 2021 · 0 repositories · arXiv:2103.07062
-
Vision Transformer for COVID-19 CXR Diagnosis using Chest X-ray Feature Corpus 12 Mar 2021 · 0 repositories · arXiv:2103.07055
-
Deepfake Video Detection Using Convolutional Vision Transformer 22 Feb 2021 · 1 repository · arXiv:2102.11126
-
Conditional Positional Encodings for Vision Transformers 22 Feb 2021 · 2 repositories · arXiv:2102.10882
-
TransReID: Transformer-based Object Re-Identification 8 Feb 2021 · 4 repositories · arXiv:2102.04378
-
PipeTransformer: Automated Elastic Pipelining for Distributed Training of Transformers 5 Feb 2021 · 1 repository · arXiv:2102.03161Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Tokens-to-Token ViT: Training Vision Transformers from Scratch on ImageNet 28 Jan 2021 · 13 repositories · arXiv:2101.11986Syntology official (archive's flag): 4 ran · 21 ran (of which 16 constructed an object rather than computing a result; 21 with no instrument failure: 1 honoured, 0 violated, 20 with no contract checked; 0 where Syntology's instrument failed) · 5 unverified (of 26 harvested samples) · 8 pointer-only (licence)
-
DAF:re: A Challenging, Crowd-Sourced, Large-Scale, Long-Tailed Dataset For Anime Character Recognition 21 Jan 2021 · 2 repositories · arXiv:2101.08674Syntology community repositories only · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Investigating the Vision Transformer Model for Image Retrieval Tasks 11 Jan 2021 · 0 repositories · arXiv:2101.03771
-
Transformer for Image Quality Assessment 30 Dec 2020 · 0 repositories · arXiv:2101.01097
-
A Survey on Visual Transformer 23 Dec 2020 · 0 repositories · arXiv:2012.12556
-
SceneFormer: Indoor Scene Generation with Transformers 17 Dec 2020 · 2 repositories · arXiv:2012.09793
-
Toward Transformer-Based Object Detection 17 Dec 2020 · 0 repositories · arXiv:2012.09958
-
AdaBins: Depth Estimation using Adaptive Bins 28 Nov 2020 · 11 repositories · arXiv:2011.14141
-
On the Effectiveness of Vision Transformers for Zero-shot Face Anti-Spoofing 16 Nov 2020 · 1 repository · arXiv:2011.08019
-
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale 22 Oct 2020 · 158 repositories · arXiv:2010.11929Syntology official: harvested, nothing ran · 307 ran (of which 165 constructed an object rather than computing a result; 286 with no instrument failure: 8 honoured, 2 violated, 276 with no contract checked; 21 where Syntology's instrument failed) · 112 unverified (of 419 harvested samples) · 154 pointer-only (licence)