{"url":"/method/vision-transformer","slug":"vision-transformer","name":"Vision Transformer","full_name":"Vision Transformer","full_name_withheld":false,"description_markdown":"The **Vision Transformer**, or **ViT**, is a model for image classification that employs a [Transformer](https://paperswithcode.com/method/transformer)-like architecture over patches of the image.  An image is split into fixed-size patches, each of them are then linearly embedded, position embeddings are added, and the resulting sequence of vectors is fed to a standard [Transformer](https://paperswithcode.com/method/transformer) encoder. In order to perform classification, the standard approach of adding an extra learnable “classification token” to the sequence is used.","description_state":"present","introduced_year":null,"introduced_by":{"title":"An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale","paper":"/paper/an-image-is-worth-16x16-words-transformers-1","first_author":"Alexey Dosovitskiy","n_authors":12,"url_abs":null,"archive_paper_url":"https://paperswithcode.com/paper/an-image-is-worth-16x16-words-transformers-1"},"source":{"url":"https://arxiv.org/abs/2010.11929v2","title":"An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale","url_on_a_paper_host":true},"code_snippet_url":"https://github.com/google-research/vision_transformer","code_snippet_url_on_a_code_host":true,"categories":[{"area":"Computer Vision","area_id":"computer-vision","collection":"Vision Transformers","url":"/methods/category/vision-transformers","pwc_aliases":["vision-transformer"]},{"area":"Computer Vision","area_id":"computer-vision","collection":"Image Models","url":"/methods/category/image-models","pwc_aliases":[]}],"n_papers_tagged":2144,"archive_num_papers":2145,"papers_newest_first":[{"paper":null,"title":"DASViT: Differentiable Architecture Search for Vision Transformer","date":"2025-07-17","arxiv_id":"2507.13079","n_code_links":0,"syntology":null},{"paper":null,"title":"Token Compression Meets Compact Vision Transformers: A Survey and Comparative Evaluation for Edge AI","date":"2025-07-13","arxiv_id":"2507.09702","n_code_links":0,"syntology":null},{"paper":null,"title":"Comparative Analysis of Vision Transformers and Traditional Deep Learning Approaches for Automated Pneumonia Detection in Chest X-Rays","date":"2025-07-11","arxiv_id":"2507.10589","n_code_links":0,"syntology":null},{"paper":"/paper/feed-forward-scenedino-for-unsupervised","title":"Feed-Forward SceneDINO for Unsupervised Semantic Scene Completion","date":"2025-07-08","arxiv_id":"2507.06230","n_code_links":1,"syntology":null},{"paper":"/paper/tile-based-vit-inference-with-visual-cluster","title":"Tile-Based ViT Inference with Visual-Cluster Priors for Zero-Shot Multi-Species Plant Identification","date":"2025-07-08","arxiv_id":"2507.06093","n_code_links":1,"syntology":null},{"paper":null,"title":"GroundingDINO-US-SAM: Text-Prompted Multi-Organ Segmentation in Ultrasound with LoRA-Tuned Vision-Language Models","date":"2025-06-30","arxiv_id":"2506.23903","n_code_links":0,"syntology":null},{"paper":"/paper/mamba-fetrack-v2-revisiting-state-space-model","title":"Mamba-FETrack V2: Revisiting State Space Model for Frame-Event based Visual Object Tracking","date":"2025-06-30","arxiv_id":"2506.23783","n_code_links":1,"syntology":null},{"paper":"/paper/attention-to-burstiness-low-rank-bilinear","title":"Attention to Burstiness: Low-Rank Bilinear Prompt Tuning","date":"2025-06-28","arxiv_id":"2506.22908","n_code_links":1,"syntology":null},{"paper":"/paper/boosting-generative-adversarial","title":"Boosting Generative Adversarial Transferability with Self-supervised Vision Transformer Features","date":"2025-06-26","arxiv_id":"2506.21046","n_code_links":1,"syntology":null},{"paper":null,"title":"Distributed Cross-Channel Hierarchical Aggregation for Foundation Models","date":"2025-06-26","arxiv_id":"2506.21411","n_code_links":0,"syntology":null},{"paper":null,"title":"X-SiT: Inherently Interpretable Surface Vision Transformers for Dementia Diagnosis","date":"2025-06-25","arxiv_id":"2506.20267","n_code_links":0,"syntology":null},{"paper":null,"title":"Vision Transformer-Based Time-Series Image Reconstruction for Cloud-Filling Applications","date":"2025-06-24","arxiv_id":"2506.19591","n_code_links":0,"syntology":null},{"paper":null,"title":"An Audio-centric Multi-task Learning Framework for Streaming Ads Targeting on Spotify","date":"2025-06-23","arxiv_id":"2506.18735","n_code_links":0,"syntology":null},{"paper":null,"title":"Deep CNN Face Matchers Inherently Support Revocable Biometric Templates","date":"2025-06-23","arxiv_id":"2506.18731","n_code_links":0,"syntology":null},{"paper":null,"title":"SSAVSV: Towards Unified Model for Self-Supervised Audio-Visual Speaker Verification","date":"2025-06-21","arxiv_id":"2506.17694","n_code_links":0,"syntology":null},{"paper":null,"title":"Exoplanet Classification through Vision Transformers with Temporal Image Analysis","date":"2025-06-19","arxiv_id":"2506.16597","n_code_links":0,"syntology":null},{"paper":null,"title":"DepthSeg: Depth prompting in remote sensing semantic segmentation","date":"2025-06-17","arxiv_id":"2506.14382","n_code_links":0,"syntology":null},{"paper":null,"title":"How Real is CARLAs Dynamic Vision Sensor? A Study on the Sim-to-Real Gap in Traffic Object Detection","date":"2025-06-16","arxiv_id":"2506.13722","n_code_links":0,"syntology":null},{"paper":null,"title":"MultiViT2: A Data-augmented Multimodal Neuroimaging Prediction Framework via Latent Diffusion Model","date":"2025-06-16","arxiv_id":"2506.13667","n_code_links":0,"syntology":null},{"paper":null,"title":"GM-LDM: Latent Diffusion Model for Brain Biomarker Identification through Functional Data-Driven Gray Matter Synthesis","date":"2025-06-15","arxiv_id":"2506.12719","n_code_links":0,"syntology":null},{"paper":"/paper/dart-differentiable-dynamic-adaptive-region","title":"DART: Differentiable Dynamic Adaptive Region Tokenizer for Vision Transformer and Mamba","date":"2025-06-12","arxiv_id":"2506.10390","n_code_links":1,"syntology":null},{"paper":null,"title":"HyBiomass: Global Hyperspectral Imagery Benchmark Dataset for Evaluating Geospatial Foundation Models in Forest Aboveground Biomass Estimation","date":"2025-06-12","arxiv_id":"2506.11314","n_code_links":0,"syntology":null},{"paper":"/paper/pipvit-patch-based-visual-interpretable","title":"PiPViT: Patch-based Visual Interpretable Prototypes for Retinal Image Analysis","date":"2025-06-12","arxiv_id":"2506.10669","n_code_links":1,"syntology":null},{"paper":null,"title":"Rethinking Random Masking in Self Distillation on ViT","date":"2025-06-12","arxiv_id":"2506.10582","n_code_links":0,"syntology":null},{"paper":"/paper/2506-10174","title":"Retrieval of Surface Solar Radiation through Implicit Albedo Recovery from Temporal Context","date":"2025-06-11","arxiv_id":"2506.10174","n_code_links":1,"syntology":null},{"paper":null,"title":"Fine-Grained Spatially Varying Material Selection in Images","date":"2025-06-10","arxiv_id":"2506.09023","n_code_links":0,"syntology":null},{"paper":null,"title":"FloorplanMAE:A self-supervised framework for complete floorplan generation from partial inputs","date":"2025-06-10","arxiv_id":"2506.08363","n_code_links":0,"syntology":null},{"paper":"/paper/patchguard-adversarially-robust-anomaly-1","title":"PatchGuard: Adversarially Robust Anomaly Detection and Localization through Vision Transformers and Pseudo Anomalies","date":"2025-06-10","arxiv_id":"2506.09237","n_code_links":2,"syntology":null},{"paper":null,"title":"Time Series Representations for Classification Lie Hidden in Pretrained Vision Transformers","date":"2025-06-10","arxiv_id":"2506.08641","n_code_links":0,"syntology":null},{"paper":null,"title":"Textile Analysis for Recycling Automation using Transfer Learning and Zero-Shot Foundation Models","date":"2025-06-06","arxiv_id":"2506.06569","n_code_links":0,"syntology":null}],"papers_shown":30,"tasks":[{"task":"/task/image-classification","name":"Image Classification","papers":322},{"task":"/task/semantic-segmentation","name":"Semantic Segmentation","papers":309},{"task":"/task/image-classification","name":"image-classification","papers":250},{"task":"/task/object-detection","name":"Object Detection","papers":197},{"task":"/task/segmentation","name":"Segmentation","papers":172},{"task":"/task/object-detection-1","name":"object-detection","papers":171},{"task":"/task/self-supervised-learning","name":"Self-Supervised Learning","papers":137},{"task":"/task/decoder","name":"Decoder","papers":125},{"task":"/task/transfer-learning","name":"Transfer Learning","papers":106},{"task":"/task/classification-1","name":"Classification","papers":101},{"task":"/task/object","name":"Object","papers":95},{"task":"/task/image-segmentation","name":"Image Segmentation","papers":78},{"task":"/task/instance-segmentation","name":"Instance Segmentation","papers":76},{"task":"/task/data-augmentation","name":"Data Augmentation","papers":74},{"task":"/task/representation-learning","name":"Representation Learning","papers":74},{"task":"/task/contrastive-learning","name":"Contrastive Learning","papers":66},{"task":"/task/diagnostic","name":"Diagnostic","papers":59},{"task":"/task/inductive-bias","name":"Inductive Bias","papers":56},{"task":"/task/retrieval","name":"Retrieval","papers":52},{"task":"/task/language-modelling","name":"Language Modelling","papers":51}],"tasks_shown":20,"n_tasks":839,"usage_by_year":[{"year":"2020","papers":7},{"year":"2021","papers":213},{"year":"2022","papers":412},{"year":"2023","papers":556},{"year":"2024","papers":644},{"year":"2025","papers":312}],"row_source":"methods_table","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/vision-transformer"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}