Methods › Computer Vision › Vision Transformers › Vision Transformer › Papers, page 12
Vision Transformer
Papers archive 2025-07-28
archive papers tagged: 2,144 · with a code link: 1,051 · where Syntology ran a sample: 328 (286 with a run with no instrument failure, 42 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (328 of 2,144 tagged: 286 with a run with no instrument failure, 42 where every run was a failure of Syntology's instrument)
Page 12 of 22: papers 1,101 to 1,200 of 2,144, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
A General Protocol to Probe Large Vision Models for 3D Physical Understanding 10 Oct 2023 · 1 repository · arXiv:2310.06836
-
A Simple and Robust Framework for Cross-Modality Medical Image Segmentation applied to Vision Transformers 9 Oct 2023 · 2 repositories · arXiv:2310.05572
-
SimPLR: A Simple and Plain Transformer for Scaling-Efficient Object Detection and Segmentation 9 Oct 2023 · 0 repositories · arXiv:2310.05920
-
Transformer Fusion with Optimal Transport 9 Oct 2023 · 1 repository · arXiv:2310.05719Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 7 unverified (of 10 harvested samples) · 10 pointer-only (licence)
-
ViTs are Everywhere: A Comprehensive Study Showcasing Vision Transformers in Different Domain 9 Oct 2023 · 0 repositories · arXiv:2310.05664
-
Low-Resolution Self-Attention for Semantic Segmentation 8 Oct 2023 · 1 repository · arXiv:2310.05026
-
FedConv: Enhancing Convolutional Neural Networks for Handling Data Heterogeneity in Federated Learning 6 Oct 2023 · 1 repository · arXiv:2310.04412Syntology official (archive's flag): 16 ran · 16 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 1 honoured, 0 violated, 9 with no contract checked; 6 where Syntology's instrument failed) · 2 unverified (of 18 harvested samples) · 4 pointer-only (licence)
-
PriViT: Vision Transformers for Fast Private Inference 6 Oct 2023 · 1 repository · arXiv:2310.04604Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
TiC: Exploring Vision Transformer in Convolution 6 Oct 2023 · 1 repository · arXiv:2310.04134
-
Exploring DINO: Emergent Properties and Limitations for Synthetic Aperture Radar Imagery 5 Oct 2023 · 0 repositories · arXiv:2310.03513
-
Beyond Random Augmentations: Pretraining with Hard Views 5 Oct 2023 · 2 repositories · arXiv:2310.03940
-
GET: Group Event Transformer for Event-Based Vision 4 Oct 2023 · 2 repositories · arXiv:2310.02642Syntology official (archive's flag): 8 ran · 8 ran (of which 6 constructed an object rather than computing a result; 8 with no instrument failure: 1 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 11 harvested samples) · 11 pointer-only (licence)
-
Improving Drumming Robot Via Attention Transformer Network 4 Oct 2023 · 0 repositories · arXiv:2310.02565
-
ObjFormer: Learning Land-Cover Changes From Paired OSM Data and Optical High-Resolution Imagery via Object-Guided Transformer 4 Oct 2023 · 1 repository · arXiv:2310.02674
-
Neural architecture impact on identifying temporally extended Reinforcement Learning tasks 4 Oct 2023 · 0 repositories · arXiv:2310.03161
-
Reinforcement Learning-based Mixture of Vision Transformers for Video Violence Recognition 4 Oct 2023 · 0 repositories · arXiv:2310.03108
-
SlowFormer: Universal Adversarial Patch for Attack on Compute and Energy Efficiency of Inference Efficient Vision Transformers 4 Oct 2023 · 1 repository · arXiv:2310.02544
-
Adapting Vision Foundation Models for Plant Phenotyping 1 Oct 2023 · 0 repositories
-
MVC: A Multi-Task Vision Transformer Network for COVID-19 Diagnosis from Chest X-ray Images 30 Sep 2023 · 0 repositories · arXiv:2310.00418
-
D³Fields: Dynamic 3D Descriptor Fields for Zero-Shot Generalizable Rearrangement 28 Sep 2023 · 0 repositories · arXiv:2309.16118
-
FLIP: Cross-domain Face Anti-spoofing with Language Guidance 28 Sep 2023 · 3 repositories · arXiv:2309.16649
-
HTC-DC Net: Monocular Height Estimation from Single Remote Sensing Images 28 Sep 2023 · 1 repository · arXiv:2309.16486
-
UVL2: A Unified Framework for Video Tampering Localization 28 Sep 2023 · 0 repositories · arXiv:2309.16126
-
Vision Transformers Need Registers 28 Sep 2023 · 6 repositories · arXiv:2309.16588Syntology official (archive's flag): 3 ran · 15 ran (of which 0 constructed an object rather than computing a result; 13 with no instrument failure: 1 honoured, 1 violated, 11 with no contract checked; 2 where Syntology's instrument failed) · 5 unverified (of 20 harvested samples) · 2 pointer-only (licence)
-
PARTICLE: Part Discovery and Contrastive Learning for Fine-grained Recognition 25 Sep 2023 · 1 repository · arXiv:2309.13822
-
A SAM-based Solution for Hierarchical Panoptic Segmentation of Crops and Weeds Competition 24 Sep 2023 · 0 repositories · arXiv:2309.13578
-
Global-correlated 3D-decoupling Transformer for Clothed Avatar Reconstruction 24 Sep 2023 · 1 repository · arXiv:2309.13524Syntology official (archive's flag): 15 ran · 15 ran (of which 0 constructed an object rather than computing a result; 14 with no instrument failure: 0 honoured, 0 violated, 14 with no contract checked; 1 where Syntology's instrument failed) · 3 unverified (of 18 harvested samples) · 18 pointer-only (licence)
-
MOSAIC: Multi-Object Segmented Arbitrary Stylization Using CLIP 24 Sep 2023 · 0 repositories · arXiv:2309.13716
-
Multi-Dimensional Hyena for Spatial Inductive Bias 24 Sep 2023 · 0 repositories · arXiv:2309.13600
-
EMGTFNet: Fuzzy Vision Transformer to decode Upperlimb sEMG signals for Hand Gestures Recognition 23 Sep 2023 · 0 repositories · arXiv:2310.03754
-
RBFormer: Improve Adversarial Robustness of Transformer by Robust Bias 23 Sep 2023 · 0 repositories · arXiv:2309.13245
-
Associative Transformer 22 Sep 2023 · 1 repository · arXiv:2309.12862
-
Masking Improves Contrastive Self-Supervised Learning for ConvNets, and Saliency Tells You Where 22 Sep 2023 · 1 repository · arXiv:2309.12757
-
Adaptive Input-image Normalization for Solving the Mode Collapse Problem in GAN-based X-ray Images 21 Sep 2023 · 0 repositories · arXiv:2309.12245
-
DAC-DETR: Divide the Attention Layers and Conquer 21 Sep 2023 · 1 repository
-
DualToken-ViT: Position-aware Efficient Vision Transformer with Dual Token Fusion 21 Sep 2023 · 0 repositories · arXiv:2309.12424
-
FLSL: Feature-level Self-supervised Learning 21 Sep 2023 · 1 repository
-
LEPARD: Learning Explicit Part Discovery for 3D Articulated Shape Reconstruction 21 Sep 2023 · 0 repositories
-
Patch n’ Pack: NaViT, a Vision Transformer for any Aspect Ratio and Resolution 21 Sep 2023 · 0 repositories
-
[Re] Masked Autoencoders Are Small Scale Vision Learners: A Reproduction Under Resource Constraints 21 Sep 2023 · 1 repository
-
[Re] On the Reproducibility of CartoonX 21 Sep 2023 · 0 repositories
-
REFINE: A Fine-Grained Medication Recommendation System Using Deep Learning and Personalized Drug Interaction Modeling 21 Sep 2023 · 0 repositories
-
Generalized Face Forgery Detection via Adaptive Learning for Pre-trained Vision Transformer 20 Sep 2023 · 1 repository · arXiv:2309.11092
-
Interpret Vision Transformers as ConvNets with Dynamic Convolutions 19 Sep 2023 · 0 repositories · arXiv:2309.10713
-
LineMarkNet: Line Landmark Detection for Valet Parking 19 Sep 2023 · 0 repositories · arXiv:2309.10475
-
Image-level supervision and self-training for transformer-based cross-modality tumor segmentation 17 Sep 2023 · 0 repositories · arXiv:2309.09246
-
MVP: Meta Visual Prompt Tuning for Few-Shot Remote Sensing Image Scene Classification 17 Sep 2023 · 0 repositories · arXiv:2309.09276
-
MMST-ViT: Climate Change-aware Crop Yield Prediction via Multi-Modal Spatial-Temporal Vision Transformer 16 Sep 2023 · 1 repository · arXiv:2309.09067Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
RingMo-lite: A Remote Sensing Multi-task Lightweight Network with CNN-Transformer Hybrid Framework 16 Sep 2023 · 0 repositories · arXiv:2309.09003
-
AnyOKP: One-Shot and Instance-Aware Object Keypoint Extraction with Pretrained ViT 15 Sep 2023 · 0 repositories · arXiv:2309.08134
-
Cross-Modal Synthesis of Structural MRI and Functional Connectivity Networks via Conditional ViT-GANs 15 Sep 2023 · 0 repositories · arXiv:2309.08160
-
Language Embedded Radiance Fields for Zero-Shot Task-Oriented Grasping 14 Sep 2023 · 0 repositories · arXiv:2309.07970
-
Virchow: A Million-Slide Digital Pathology Foundation Model 14 Sep 2023 · 1 repository · arXiv:2309.07778Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 9 harvested samples) · 9 pointer-only (licence)
-
A 3M-Hybrid Model for the Restoration of Unique Giant Murals: A Case Study on the Murals of Yongle Palace 12 Sep 2023 · 0 repositories · arXiv:2309.06194
-
Feature Aggregation Network for Building Extraction from High-resolution Remote Sensing Images 12 Sep 2023 · 0 repositories · arXiv:2309.06017
-
Toward a Deeper Understanding: RetNet Viewed through Convolution 11 Sep 2023 · 1 repository · arXiv:2309.05375
-
Restoring Snow-Degraded Single Images With Wavelet in Vision Transformer 11 Sep 2023 · 2 repositories
-
DeViT: Decomposing Vision Transformers for Collaborative Inference in Edge Devices 10 Sep 2023 · 0 repositories · arXiv:2309.05015
-
How to Evaluate Semantic Communications for Images with ViTScore Metric? 9 Sep 2023 · 0 repositories · arXiv:2309.04891
-
Leveraging Pretrained Image-text Models for Improving Audio-Visual Learning 8 Sep 2023 · 0 repositories · arXiv:2309.04628
-
Adapting Self-Supervised Representations to Multi-Domain Setups 7 Sep 2023 · 0 repositories · arXiv:2309.03999
-
S-Adapter: Generalizing Vision Transformer for Face Anti-Spoofing with Statistical Tokens 7 Sep 2023 · 3 repositories · arXiv:2309.04038
-
Improving diagnosis and prognosis of lung cancer using vision transformers: A scoping review 6 Sep 2023 · 0 repositories · arXiv:2309.02783
-
A survey on efficient vision transformers: algorithms, techniques, and performance benchmarking 5 Sep 2023 · 0 repositories · arXiv:2309.02031
-
Compressing Vision Transformers for Low-Resource Visual Learning 5 Sep 2023 · 1 repository · arXiv:2309.02617
-
Domain Adaptation for Efficiently Fine-tuning Vision Transformer with Encrypted Images 5 Sep 2023 · 0 repositories · arXiv:2309.02556
-
ExMobileViT: Lightweight Classifier Extension for Mobile Vision Transformer 4 Sep 2023 · 0 repositories · arXiv:2309.01310
-
Locality-Aware Hyperspectral Classification 4 Sep 2023 · 1 repository · arXiv:2309.01561
-
Semantic-Constraint Matching Transformer for Weakly Supervised Object Localization 4 Sep 2023 · 0 repositories · arXiv:2309.01331
-
Contrastive Feature Masking Open-Vocabulary Vision Transformer 2 Sep 2023 · 0 repositories · arXiv:2309.00775
-
Beyond Self-Attention: Deformable Large Kernel Attention for Medical Image Segmentation 31 Aug 2023 · 1 repository · arXiv:2309.00121
-
Emergence of Segmentation with Minimalistic White-Box Transformers 30 Aug 2023 · 1 repository · arXiv:2308.16271Syntology official (archive's flag): 2 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
1st Place Solution for the 5th LSVOS Challenge: Video Instance Segmentation 28 Aug 2023 · 1 repository · arXiv:2308.14392
-
Fast Feedforward Networks 28 Aug 2023 · 4 repositories · arXiv:2308.14711
-
FIRE: Food Image to REcipe generation 28 Aug 2023 · 1 repository · arXiv:2308.14391Syntology official (archive's flag): 12 ran · 12 ran (of which 2 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 2 where Syntology's instrument failed) · 4 unverified (of 16 harvested samples) · 16 pointer-only (licence)
-
VideoCutLER: Surprisingly Simple Unsupervised Video Instance Segmentation 28 Aug 2023 · 1 repository · arXiv:2308.14710Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
A comprehensive review on Plant Leaf Disease detection using Deep learning 27 Aug 2023 · 0 repositories · arXiv:2308.14087
-
DETDet: Dual Ensemble Teeth Detection 27 Aug 2023 · 1 repository · arXiv:2308.14070
-
Gaze-Informed Vision Transformers: Predicting Driving Decisions Under Uncertainty 26 Aug 2023 · 1 repository · arXiv:2308.13969
-
A Re-Parameterized Vision Transformer (ReVT) for Domain-Generalized Semantic Segmentation 25 Aug 2023 · 1 repository · arXiv:2308.13331
-
ACC-UNet: A Completely Convolutional UNet model for the 2020s 25 Aug 2023 · 1 repository · arXiv:2308.13680
-
An investigation into the impact of deep learning model choice on sex and race bias in cardiac MR segmentation 25 Aug 2023 · 0 repositories · arXiv:2308.13415
-
Linear Oscillation: A Novel Activation Function for Vision Transformer 25 Aug 2023 · 0 repositories · arXiv:2308.13670
-
Full-dose Whole-body PET Synthesis from Low-dose PET Using High-efficiency Denoising Diffusion Probabilistic Model: PET Consistency Model 24 Aug 2023 · 1 repository · arXiv:2308.13072
-
Towards Hierarchical Regional Transformer-based Multiple Instance Learning 24 Aug 2023 · 0 repositories · arXiv:2308.12634
-
Local Distortion Aware Efficient Transformer Adaptation for Image Quality Assessment 23 Aug 2023 · 0 repositories · arXiv:2308.12001
-
Masking Strategies for Background Bias Removal in Computer Vision Models 23 Aug 2023 · 1 repository · arXiv:2308.12127
-
MOFO: MOtion FOcused Self-Supervision for Video Understanding 23 Aug 2023 · 1 repository · arXiv:2308.12447Syntology official (archive's flag): 1 ran · 1 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified; the one sample that ran constructed an object rather than computing a result (of 1 harvested sample) · 1 pointer-only (licence)
-
Vision Transformer Adapters for Generalizable Multitask Learning 23 Aug 2023 · 0 repositories · arXiv:2308.12372
-
Deep learning for automated materials characterisation in core-loss electron energy loss spectroscopy 22 Aug 2023 · 1 repository
-
TurboViT: Generating Fast Vision Transformers via Generative Architecture Search 22 Aug 2023 · 0 repositories · arXiv:2308.11421
-
Joint learning of images and videos with a single Vision Transformer 21 Aug 2023 · 0 repositories · arXiv:2308.10533
-
Vision Transformer Pruning Via Matrix Decomposition 21 Aug 2023 · 0 repositories · arXiv:2308.10839
-
FedSIS: Federated Split Learning with Intermediate Representation Sampling for Privacy-preserving Generalized Face Presentation Attack Detection 20 Aug 2023 · 2 repositories · arXiv:2308.10236
-
Towards a High-Performance Object Detector: Insights from Drone Detection Using ViT and CNN-based Deep Learning Models 19 Aug 2023 · 0 repositories · arXiv:2308.09899
-
SimFIR: A Simple Framework for Fisheye Image Rectification with Self-supervised Representation Learning 17 Aug 2023 · 0 repositories · arXiv:2308.09040
-
DSAT-Net: Dual Spatial Attention Transformer for Building Extraction from Aerial Images 16 Aug 2023 · 1 repository
-
SkinDistilViT: Lightweight Vision Transformer for Skin Lesion Classification 16 Aug 2023 · 1 repository · arXiv:2308.08669
-
Enhancing Network Initialization for Medical AI Models Using Large-Scale, Unlabeled Natural Images 15 Aug 2023 · 2 repositories · arXiv:2308.07688
-
Fast Machine Unlearning Without Retraining Through Selective Synaptic Dampening 15 Aug 2023 · 1 repository · arXiv:2308.07707Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 6 harvested samples) · 5 pointer-only (licence)