Methods › Computer Vision › Vision Transformers › Vision Transformer
Vision Transformer
Introduced by Alexey Dosovitskiy et al. in An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
archive 2025-07-28 Description, source and code snippet are the archive's method entry.
The Vision Transformer, or ViT, is a model for image classification that employs a Transformer-like architecture over patches of the image. An image is split into fixed-size patches, each of them are then linearly embedded, position embeddings are added, and the resulting sequence of vectors is fed to a standard Transformer encoder. In order to perform classification, the standard approach of adding an extra learnable “classification token” to the sequence is used.
Papers archive 2025-07-28
30 shown of 2,144, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.
-
DASViT: Differentiable Architecture Search for Vision Transformer 17 Jul 2025 · 0 repositories · arXiv:2507.13079
-
Token Compression Meets Compact Vision Transformers: A Survey and Comparative Evaluation for Edge AI 13 Jul 2025 · 0 repositories · arXiv:2507.09702
-
Comparative Analysis of Vision Transformers and Traditional Deep Learning Approaches for Automated Pneumonia Detection in Chest X-Rays 11 Jul 2025 · 0 repositories · arXiv:2507.10589
-
Feed-Forward SceneDINO for Unsupervised Semantic Scene Completion 8 Jul 2025 · 1 repository · arXiv:2507.06230
-
Tile-Based ViT Inference with Visual-Cluster Priors for Zero-Shot Multi-Species Plant Identification 8 Jul 2025 · 1 repository · arXiv:2507.06093
-
GroundingDINO-US-SAM: Text-Prompted Multi-Organ Segmentation in Ultrasound with LoRA-Tuned Vision-Language Models 30 Jun 2025 · 0 repositories · arXiv:2506.23903
-
Mamba-FETrack V2: Revisiting State Space Model for Frame-Event based Visual Object Tracking 30 Jun 2025 · 1 repository · arXiv:2506.23783
-
Attention to Burstiness: Low-Rank Bilinear Prompt Tuning 28 Jun 2025 · 1 repository · arXiv:2506.22908
-
Boosting Generative Adversarial Transferability with Self-supervised Vision Transformer Features 26 Jun 2025 · 1 repository · arXiv:2506.21046
-
Distributed Cross-Channel Hierarchical Aggregation for Foundation Models 26 Jun 2025 · 0 repositories · arXiv:2506.21411
-
X-SiT: Inherently Interpretable Surface Vision Transformers for Dementia Diagnosis 25 Jun 2025 · 0 repositories · arXiv:2506.20267
-
Vision Transformer-Based Time-Series Image Reconstruction for Cloud-Filling Applications 24 Jun 2025 · 0 repositories · arXiv:2506.19591
-
An Audio-centric Multi-task Learning Framework for Streaming Ads Targeting on Spotify 23 Jun 2025 · 0 repositories · arXiv:2506.18735
-
Deep CNN Face Matchers Inherently Support Revocable Biometric Templates 23 Jun 2025 · 0 repositories · arXiv:2506.18731
-
SSAVSV: Towards Unified Model for Self-Supervised Audio-Visual Speaker Verification 21 Jun 2025 · 0 repositories · arXiv:2506.17694
-
Exoplanet Classification through Vision Transformers with Temporal Image Analysis 19 Jun 2025 · 0 repositories · arXiv:2506.16597
-
DepthSeg: Depth prompting in remote sensing semantic segmentation 17 Jun 2025 · 0 repositories · arXiv:2506.14382
-
How Real is CARLAs Dynamic Vision Sensor? A Study on the Sim-to-Real Gap in Traffic Object Detection 16 Jun 2025 · 0 repositories · arXiv:2506.13722
-
MultiViT2: A Data-augmented Multimodal Neuroimaging Prediction Framework via Latent Diffusion Model 16 Jun 2025 · 0 repositories · arXiv:2506.13667
-
GM-LDM: Latent Diffusion Model for Brain Biomarker Identification through Functional Data-Driven Gray Matter Synthesis 15 Jun 2025 · 0 repositories · arXiv:2506.12719
-
DART: Differentiable Dynamic Adaptive Region Tokenizer for Vision Transformer and Mamba 12 Jun 2025 · 1 repository · arXiv:2506.10390
-
HyBiomass: Global Hyperspectral Imagery Benchmark Dataset for Evaluating Geospatial Foundation Models in Forest Aboveground Biomass Estimation 12 Jun 2025 · 0 repositories · arXiv:2506.11314
-
PiPViT: Patch-based Visual Interpretable Prototypes for Retinal Image Analysis 12 Jun 2025 · 1 repository · arXiv:2506.10669
-
Rethinking Random Masking in Self Distillation on ViT 12 Jun 2025 · 0 repositories · arXiv:2506.10582
-
Retrieval of Surface Solar Radiation through Implicit Albedo Recovery from Temporal Context 11 Jun 2025 · 1 repository · arXiv:2506.10174
-
Fine-Grained Spatially Varying Material Selection in Images 10 Jun 2025 · 0 repositories · arXiv:2506.09023
-
FloorplanMAE:A self-supervised framework for complete floorplan generation from partial inputs 10 Jun 2025 · 0 repositories · arXiv:2506.08363
-
PatchGuard: Adversarially Robust Anomaly Detection and Localization through Vision Transformers and Pseudo Anomalies 10 Jun 2025 · 2 repositories · arXiv:2506.09237
-
Time Series Representations for Classification Lie Hidden in Pretrained Vision Transformers 10 Jun 2025 · 0 repositories · arXiv:2506.08641
-
Textile Analysis for Recycling Automation using Transfer Learning and Zero-Shot Foundation Models 6 Jun 2025 · 0 repositories · arXiv:2506.06569
Tasks archive 2025-07-28
20 shown of 839 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.
Usage over time archive 2025-07-28
Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).
Categories archive 2025-07-28
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections