Methods › Computer Vision › Vision Transformers › Vision Transformer › Papers, page 21
Vision Transformer
Papers archive 2025-07-28
archive papers tagged: 2,144 · with a code link: 1,051 · where Syntology ran a sample: 328 (286 with a run with no instrument failure, 42 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (328 of 2,144 tagged: 286 with a run with no instrument failure, 42 where every run was a failure of Syntology's instrument)
Page 21 of 22: papers 2,001 to 2,100 of 2,144, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Adversarial Robustness Comparison of Vision Transformer and MLP-Mixer to CNNs 6 Oct 2021 · 1 repository · arXiv:2110.02797
-
MobileViT: Light-weight, General-purpose, and Mobile-friendly Vision Transformer 5 Oct 2021 · 31 repositories · arXiv:2110.02178Syntology official: no sample here; runs from other or unrecorded repositories · 53 ran (of which 31 constructed an object rather than computing a result; 44 with no instrument failure: 1 honoured, 0 violated, 43 with no contract checked; 9 where Syntology's instrument failed) · 15 unverified (of 68 harvested samples) · 18 pointer-only (licence)
-
A free lunch from ViT:Adaptive Attention Multi-scale Fusion Transformer for Fine-grained Visual Recognition 4 Oct 2021 · 0 repositories · arXiv:2110.01240
-
VTAMIQ: Transformers for Attention Modulated Image Quality Assessment 4 Oct 2021 · 1 repository · arXiv:2110.01655
-
Implicit and Explicit Attention for Zero-Shot Learning 2 Oct 2021 · 1 repository · arXiv:2110.00860
-
Learning to Predict Trustworthiness with Steep Slope Loss 30 Sep 2021 · 1 repository · arXiv:2110.00054Syntology official (archive's flag): 1 ran · 1 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified; the one sample that ran constructed an object rather than computing a result (of 1 harvested sample)
-
Are Vision Transformers Robust to Patch-wise Perturbations? 29 Sep 2021 · 0 repositories
-
CCTrans: Simplifying and Improving Crowd Counting with Transformer 29 Sep 2021 · 2 repositories · arXiv:2109.14483Syntology 5 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 6 harvested samples) · 2 pointer-only (licence)
-
HFSP: A Hardware-friendly Soft Pruning Framework for Vision Transformers 29 Sep 2021 · 0 repositories
-
Localizing Objects with Self-Supervised Transformers and no Labels 29 Sep 2021 · 2 repositories · arXiv:2109.14279
-
Privacy-preserving Task-Agnostic Vision Transformer for Image Processing 29 Sep 2021 · 1 repository
-
Test Time Robustification of Deep Models via Adaptation and Augmentation 29 Sep 2021 · 0 repositories
-
UFO-ViT: High Performance Linear Vision Transformer without Softmax 29 Sep 2021 · 1 repository · arXiv:2109.14382
-
Fine-tuning Vision Transformers for the Prediction of State Variables in Ising Models 28 Sep 2021 · 0 repositories · arXiv:2109.13925
-
PASS: An ImageNet replacement for self-supervised pretraining without humans 27 Sep 2021 · 1 repository · arXiv:2109.13228Syntology official: no sample here; runs from other or unrecorded repositories · 11 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 1 honoured, 0 violated, 9 with no contract checked; 1 where Syntology's instrument failed) · 3 unverified (of 14 harvested samples) · 9 pointer-only (licence)
-
Vision Transformer Hashing for Image Retrieval 26 Sep 2021 · 1 repository · arXiv:2109.12564
-
ViT Cane: Visual Assistant for the Visually Impaired 26 Sep 2021 · 0 repositories · arXiv:2109.13857
-
BiTr-Unet: a CNN-Transformer Combined Network for MRI Brain Tumor Segmentation 25 Sep 2021 · 1 repository · arXiv:2109.12271
-
TEMGNet: Deep Transformer-based Decoding of Upperlimb sEMG for Hand Gestures Recognition 25 Sep 2021 · 0 repositories · arXiv:2109.12379
-
Improving 360 Monocular Depth Estimation via Non-local Dense Prediction Transformer and Joint Supervised and Self-supervised Learning 22 Sep 2021 · 1 repository · arXiv:2109.10563Syntology 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 1 pointer-only (licence)
-
DS-Net++: Dynamic Weight Slicing for Efficient Inference in CNNs and Transformers 21 Sep 2021 · 1 repository · arXiv:2109.10060
-
MFEViT: A Robust Lightweight Transformer-based Network for Multimodal 2D+3D Facial Expression Recognition 20 Sep 2021 · 0 repositories · arXiv:2109.13086
-
UNetFormer: A UNet-like Transformer for Efficient Semantic Segmentation of Remote Sensing Urban Scene Imagery 18 Sep 2021 · 1 repository · arXiv:2109.08937Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 2 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Complementary Feature Enhanced Network with Vision Transformer for Image Dehazing 15 Sep 2021 · 1 repository · arXiv:2109.07100
-
Vision Transformer for Learning Driving Policies in Complex Multi-Agent Environments 14 Sep 2021 · 0 repositories · arXiv:2109.06514
-
Improving Video-Text Retrieval by Multi-Stream Corpus Alignment and Dual Softmax Loss 9 Sep 2021 · 2 repositories · arXiv:2109.04290
-
Vision Transformers For Weeds and Crops Classification Of High Resolution UAV Images 6 Sep 2021 · 0 repositories · arXiv:2109.02716
-
A Battle of Network Structures: An Empirical Study of CNN, Transformer, and MLP 30 Aug 2021 · 1 repository · arXiv:2108.13002
-
Exploring and Improving Mobile Level Vision Transformers 30 Aug 2021 · 0 repositories · arXiv:2108.13015
-
Towards Fine-grained Image Classification with Generative Adversarial Networks and Facial Landmark Detection 28 Aug 2021 · 1 repository · arXiv:2109.00891
-
Transformer for Single Image Super-Resolution 25 Aug 2021 · 1 repository · arXiv:2108.11084Syntology official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 10 harvested samples) · 4 pointer-only (licence)
-
ZS-SLR: Zero-Shot Sign Language Recognition from RGB-D Videos 23 Aug 2021 · 0 repositories · arXiv:2108.10059
-
Construction material classification on imbalanced datasets using Vision Transformer (ViT) architecture 21 Aug 2021 · 0 repositories · arXiv:2108.09527
-
Convolutional Neural Network (CNN) vs Vision Transformer (ViT) for Digital Holography 20 Aug 2021 · 0 repositories · arXiv:2108.09147
-
Causal Attention for Unbiased Visual Recognition 19 Aug 2021 · 1 repository · arXiv:2108.08782Syntology official (archive's flag): 2 ran · 2 ran (of which 2 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified; every one of the 2 samples that ran constructed an object rather than computing a result (of 2 harvested samples) · 2 pointer-only (licence)
-
Boosting Salient Object Detection with Transformer-based Asymmetric Bilateral U-Net 17 Aug 2021 · 1 repository · arXiv:2108.07851
-
TVT: Transferable Vision Transformer for Unsupervised Domain Adaptation 12 Aug 2021 · 1 repository · arXiv:2108.05988Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 2 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
RaftMLP: How Much Can Be Done Without Attention and with Less Spatial Locality? 9 Aug 2021 · 2 repositories · arXiv:2108.04384
-
Vision Transformer for femur fracture classification 7 Aug 2021 · 0 repositories · arXiv:2108.03414
-
Token Shift Transformer for Video Classification 5 Aug 2021 · 3 repositories · arXiv:2108.02432Syntology community repositories only · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Dynamic Feature Regularized Loss for Weakly Supervised Semantic Segmentation 3 Aug 2021 · 0 repositories · arXiv:2108.01296
-
Vision Transformer with Progressive Sampling 3 Aug 2021 · 1 repository · arXiv:2108.01684Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 7 harvested samples) · 2 pointer-only (licence)
-
Congested Crowd Instance Localization with Dilated Convolutional Swin Transformer 2 Aug 2021 · 1 repository · arXiv:2108.00584
-
Multi-Head Self-Attention via Vision Transformer for Zero-Shot Learning 30 Jul 2021 · 2 repositories · arXiv:2108.00045
-
Go Wider Instead of Deeper 25 Jul 2021 · 1 repository · arXiv:2107.11817
-
Weakly Supervised Global-Local Feature Learning for Cervical Cytology Image Analysis 20 Jul 2021 · 0 repositories
-
RAMS-Trans: Recurrent Attention Multi-scale Transformer forFine-grained Image Recognition 17 Jul 2021 · 0 repositories · arXiv:2107.08192
-
GLiT: Neural Architecture Search for Global and Local Image Transformer 7 Jul 2021 · 2 repositories · arXiv:2107.02960Syntology community repositories only · 2 ran (of which 2 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified; every one of the 2 samples that ran constructed an object rather than computing a result (of 4 harvested samples)
-
Learning Vision Transformer with Squeeze and Excitation for Facial Expression Recognition 7 Jul 2021 · 0 repositories · arXiv:2107.03107
-
Scopeformer: n-CNN-ViT Hybrid Model for Intracranial Hemorrhage Classification 7 Jul 2021 · 0 repositories · arXiv:2107.04575
-
Feature Fusion Vision Transformer for Fine-Grained Visual Categorization 6 Jul 2021 · 1 repository · arXiv:2107.02341
-
Vision Xformers: Efficient Attention for Image Classification 5 Jul 2021 · 2 repositories · arXiv:2107.02239
-
What Makes for Hierarchical Vision Transformer? 5 Jul 2021 · 0 repositories · arXiv:2107.02174
-
COVID-VIT: Classification of COVID-19 from CT chest images based on vision transformer models 4 Jul 2021 · 1 repository · arXiv:2107.01682
-
Learning Efficient Vision Transformers via Fine-Grained Manifold Distillation 3 Jul 2021 · 1 repository · arXiv:2107.01378
-
AutoFormer: Searching Transformers for Visual Recognition 1 Jul 2021 · 2 repositories · arXiv:2107.00651
-
Focal Self-attention for Local-Global Interactions in Vision Transformers 1 Jul 2021 · 3 repositories · arXiv:2107.00641Syntology official (archive's flag): 1 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 2 pointer-only (licence)
-
Looking Outside the Window: Wide-Context Transformer for the Semantic Segmentation of High-Resolution Remote Sensing Images 29 Jun 2021 · 1 repository · arXiv:2106.15754
-
Multi-Exit Vision Transformer for Dynamic Inference 29 Jun 2021 · 0 repositories · arXiv:2106.15183
-
Rethinking Token-Mixing MLP for MLP-based Vision Backbone 28 Jun 2021 · 0 repositories · arXiv:2106.14882
-
OffRoadTranSeg: Semi-Supervised Segmentation using Transformers on OffRoad environments 26 Jun 2021 · 0 repositories · arXiv:2106.13963
-
PVT v2: Improved Baselines with Pyramid Vision Transformer 25 Jun 2021 · 18 repositories · arXiv:2106.13797
-
Exploring Corruption Robustness: Inductive Biases in Vision Transformers and MLP-Mixers 24 Jun 2021 · 1 repository · arXiv:2106.13122
-
IA-RED²: Interpretability-Aware Redundancy Reduction for Vision Transformers 23 Jun 2021 · 0 repositories · arXiv:2106.12620
-
Instance-based Vision Transformer for Subtyping of Papillary Renal Cell Carcinoma in Histopathological Image 23 Jun 2021 · 1 repository · arXiv:2106.12265
-
P2T: Pyramid Pooling Transformer for Scene Understanding 22 Jun 2021 · 4 repositories · arXiv:2106.12011
-
Exploring Vision Transformers for Fine-grained Classification 19 Jun 2021 · 1 repository · arXiv:2106.10587
-
Video Super-Resolution Transformer 12 Jun 2021 · 1 repository · arXiv:2106.06847
-
MlTr: Multi-label Classification with Transformer 11 Jun 2021 · 1 repository · arXiv:2106.06195
-
ViT-Inception-GAN for Image Colourising 11 Jun 2021 · 0 repositories · arXiv:2106.06321
-
MST: Masked Self-Supervised Transformer for Visual Representation 10 Jun 2021 · 0 repositories · arXiv:2106.05656
-
Scaling Vision with Sparse Mixture of Experts 10 Jun 2021 · 1 repository · arXiv:2106.05974Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Grounding inductive biases in natural images:invariance stems from variations in data 9 Jun 2021 · 1 repository · arXiv:2106.05121Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Towards Training Stronger Video Vision Transformers for EPIC-KITCHENS-100 Action Recognition 9 Jun 2021 · 1 repository · arXiv:2106.05058
-
On the Connection between Local Attention and Dynamic Depth-wise Convolution 8 Jun 2021 · 1 repository · arXiv:2106.04263Syntology official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 11 with no instrument failure: 1 honoured, 1 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 5 unverified (of 16 harvested samples) · 9 pointer-only (licence)
-
MVT: Mask Vision Transformer for Facial Expression Recognition in the wild 8 Jun 2021 · 0 repositories · arXiv:2106.04520
-
Scaling Vision Transformers 8 Jun 2021 · 1 repository · arXiv:2106.04560
-
Person Re-Identification with a Locally Aware Transformer 7 Jun 2021 · 1 repository · arXiv:2106.03720Syntology official (archive's flag): 2 ran · 2 ran (of which 2 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified; every one of the 2 samples that ran constructed an object rather than computing a result (of 2 harvested samples)
-
Reveal of Vision Transformers Robustness against Adversarial Attacks 7 Jun 2021 · 0 repositories · arXiv:2106.03734
-
ViTAE: Vision Transformer Advanced by Exploring Intrinsic Inductive Bias 7 Jun 2021 · 2 repositories · arXiv:2106.03348Syntology official: no sample here; runs from other or unrecorded repositories · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Vision Transformers with Hierarchical Attention 6 Jun 2021 · 3 repositories · arXiv:2106.03180
-
Container: Context Aggregation Network 2 Jun 2021 · 4 repositories · arXiv:2106.01401Syntology official (archive's flag): 1 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 7 harvested samples)
-
TransVOS: Video Object Segmentation with Transformers 1 Jun 2021 · 1 repository · arXiv:2106.00588
-
You Only Look at One Sequence: Rethinking Transformer in Vision through Object Detection 1 Jun 2021 · 2 repositories · arXiv:2106.00666Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 1 honoured, 0 violated, 1 with no contract checked; 3 where Syntology's instrument failed) · 5 unverified (of 10 harvested samples) · 5 pointer-only (licence)
-
Analogous to Evolutionary Algorithm: Designing a Unified Sequence Model 31 May 2021 · 1 repository · arXiv:2105.15089
-
Gaze Estimation using Transformer 30 May 2021 · 1 repository · arXiv:2105.14424
-
TransMatcher: Deep Image Matching Through Transformers for Generalizable Person Re-identification 30 May 2021 · 2 repositories · arXiv:2105.14432Syntology official (archive's flag): 5 ran · 13 ran (of which 4 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 4 where Syntology's instrument failed) · 8 unverified (of 21 harvested samples) · 1 pointer-only (licence)
-
FoveaTer: Foveated Transformer for Image Classification 29 May 2021 · 0 repositories · arXiv:2105.14173
-
Less is More: Pay Less Attention in Vision Transformers 29 May 2021 · 2 repositories · arXiv:2105.14217
-
KVT: k-NN Attention for Boosting Vision Transformers 28 May 2021 · 1 repository · arXiv:2106.00515Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples)
-
Nested Hierarchical Transformer: Towards Accurate, Data-Efficient and Interpretable Visual Understanding 26 May 2021 · 6 repositories · arXiv:2105.12723Syntology official (archive's flag): 9 ran · 19 ran (of which 0 constructed an object rather than computing a result; 18 with no instrument failure: 1 honoured, 1 violated, 16 with no contract checked; 1 where Syntology's instrument failed) · 7 unverified (of 26 harvested samples) · 6 pointer-only (licence)
-
Federated Split Task-Agnostic Vision Transformer for COVID-19 CXR Diagnosis 21 May 2021 · 0 repositories
-
Grounding inductive biases in natural images: invariance stems from variations in data 21 May 2021 · 1 repository
-
Single-Layer Vision Transformers for More Accurate Early Exits with Less Overhead 19 May 2021 · 0 repositories · arXiv:2105.09121
-
Vision Transformer for Fast and Efficient Scene Text Recognition 18 May 2021 · 3 repositories · arXiv:2105.08582
-
Towards Robust Vision Transformer 17 May 2021 · 2 repositories · arXiv:2105.07926
-
Vision Transformers are Robust Learners 17 May 2021 · 1 repository · arXiv:2105.07581Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 1 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 1 pointer-only (licence)
-
Are Convolutional Neural Networks or Transformers more like human vision? 15 May 2021 · 1 repository · arXiv:2105.07197
-
Manipulation Detection in Satellite Images Using Vision Transformer 13 May 2021 · 0 repositories · arXiv:2105.06373
-
A Large-Scale Benchmark for Food Image Segmentation 12 May 2021 · 2 repositories · arXiv:2105.05409Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples) · 1 pointer-only (licence)