Methods › Computer Vision › Vision Transformers › Vision Transformer › Papers, page 14
Vision Transformer
Papers archive 2025-07-28
archive papers tagged: 2,144 · with a code link: 1,051 · where Syntology ran a sample: 328 (286 with a run with no instrument failure, 42 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (328 of 2,144 tagged: 286 with a run with no instrument failure, 42 where every run was a failure of Syntology's instrument)
Page 14 of 22: papers 1,301 to 1,400 of 2,144, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
nnMobileNet: Rethinking CNN for Retinopathy Research 2 Jun 2023 · 2 repositories · arXiv:2306.01289
-
Auto-Spikformer: Spikformer Architecture Search 1 Jun 2023 · 0 repositories · arXiv:2306.00807
-
Hiera: A Hierarchical Vision Transformer without the Bells-and-Whistles 1 Jun 2023 · 4 repositories · arXiv:2306.00989Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 6 harvested samples)
-
Lightweight Vision Transformer with Bidirectional Interaction 1 Jun 2023 · 1 repository · arXiv:2306.00396Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Diagnosis and Prognosis of Head and Neck Cancer Patients using Artificial Intelligence 31 May 2023 · 0 repositories · arXiv:2306.00034
-
Humans in 4D: Reconstructing and Tracking Humans with Transformers 31 May 2023 · 1 repository · arXiv:2305.20091Syntology official (archive's flag): 4 ran · 4 ran (of which 1 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples)
-
LOWA: Localize Objects in the Wild with Attributes 31 May 2023 · 0 repositories · arXiv:2305.20047
-
Prompt-Based Tuning of Transformer Models for Multi-Center Medical Image Segmentation of Head and Neck Cancer 30 May 2023 · 0 repositories · arXiv:2305.18948
-
Vision Transformers for Mobile Applications: A Short Survey 30 May 2023 · 0 repositories · arXiv:2305.19365
-
Solar Irradiance Anticipative Transformer 29 May 2023 · 1 repository · arXiv:2305.18487
-
Reconstructing Sea Surface Temperature Images: A Masked Autoencoder Approach for Cloud Masking and Reconstruction 28 May 2023 · 0 repositories · arXiv:2306.00835
-
Zero-TPrune: Zero-Shot Token Pruning through Leveraging of the Attention Graph in Pre-Trained Transformers 27 May 2023 · 0 repositories · arXiv:2305.17328
-
COMCAT: Towards Efficient Compression and Customization of Attention-Based Vision Models 26 May 2023 · 1 repository · arXiv:2305.17235Syntology official (archive's flag): 1 ran · 1 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified; the one sample that ran constructed an object rather than computing a result (of 2 harvested samples) · 2 pointer-only (licence)
-
Do We Really Need a Large Number of Visual Prompts? 26 May 2023 · 0 repositories · arXiv:2305.17223
-
GenerateCT: Text-Conditional Generation of 3D Chest CT Volumes 25 May 2023 · 1 repository · arXiv:2305.16037Syntology official (archive's flag): 14 ran · 14 ran (of which 0 constructed an object rather than computing a result; 13 with no instrument failure: 1 honoured, 4 violated, 8 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 15 harvested samples) · 5 pointer-only (licence)
-
Multi-scale Efficient Graph-Transformer for Whole Slide Image Classification 25 May 2023 · 0 repositories · arXiv:2305.15773
-
Learning UI-to-Code Reverse Generator Using Visual Critic Without Rendering 24 May 2023 · 0 repositories · arXiv:2305.14637
-
Weakly Supervised 3D Open-vocabulary Segmentation 23 May 2023 · 1 repository · arXiv:2305.14093Syntology official (archive's flag): 12 ran · 12 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 2 where Syntology's instrument failed) · 2 unverified (of 14 harvested samples) · 10 pointer-only (licence)
-
DeepJSCC-l++: Robust and Bandwidth-Adaptive Wireless Image Transmission 22 May 2023 · 1 repository · arXiv:2305.13161
-
Efficient Large-Scale Visual Representation Learning And Evaluation 22 May 2023 · 0 repositories · arXiv:2305.13399
-
Materialistic: Selecting Similar Materials in Images 22 May 2023 · 0 repositories · arXiv:2305.13291
-
Spatiotemporal Attention-based Semantic Compression for Real-time Video Recognition 22 May 2023 · 0 repositories · arXiv:2305.12796
-
TSPTQ-ViT: Two-scaled post-training quantization for vision transformer 22 May 2023 · 0 repositories · arXiv:2305.12901
-
Why current rain denoising models fail on CycleGAN created rain images in autonomous driving 22 May 2023 · 0 repositories · arXiv:2305.12983
-
Comparative Analysis of Deep Learning Models for Brand Logo Classification in Real-World Scenarios 20 May 2023 · 0 repositories · arXiv:2305.12242
-
Multimodal Web Navigation with Instruction-Finetuned Foundation Models 19 May 2023 · 0 repositories · arXiv:2305.11854
-
Surgical-VQLA: Transformer with Gated Vision-Language Embedding for Visual Question Localized-Answering in Robotic Surgery 19 May 2023 · 2 repositories · arXiv:2305.11692Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 7 harvested samples) · 4 pointer-only (licence)
-
Boost Vision Transformer with GPU-Friendly Sparsity and Quantization 18 May 2023 · 0 repositories · arXiv:2305.10727
-
A survey of the Vision Transformers and their CNN-Transformer based Variants 17 May 2023 · 0 repositories · arXiv:2305.09880
-
Blind Image Quality Assessment via Transformer Predicted Error Map and Perceptual Quality Token 16 May 2023 · 1 repository · arXiv:2305.09353
-
CB-HVTNet: A channel-boosted hybrid vision transformer network for lymphocyte assessment in histopathological images 16 May 2023 · 0 repositories · arXiv:2305.09211
-
AutoRecon: Automated 3D Object Discovery and Reconstruction 15 May 2023 · 0 repositories · arXiv:2305.08810
-
MaxViT-UNet: Multi-Axis Attention for Medical Image Segmentation 15 May 2023 · 2 repositories · arXiv:2305.08396
-
Meta-Polyp: a baseline for efficient Polyp segmentation 13 May 2023 · 2 repositories · arXiv:2305.07848
-
A Survey on Segment Anything Model (SAM): Vision Foundation Model Meets Prompt Engineering 12 May 2023 · 0 repositories · arXiv:2306.06211
-
Hausdorff Distance Matching with Adaptive Query Denoising for Rotated Detection Transformer 12 May 2023 · 1 repository · arXiv:2305.07598
-
ViT Unified: Joint Fingerprint Recognition and Presentation Attack Detection 12 May 2023 · 0 repositories · arXiv:2305.07602
-
Salient Mask-Guided Vision Transformer for Fine-Grained Classification 11 May 2023 · 1 repository · arXiv:2305.07102
-
Undercover Deepfakes: Detecting Fake Segments in Videos 11 May 2023 · 2 repositories · arXiv:2305.06564
-
WeLayout: WeChat Layout Analysis System for the ICDAR 2023 Competition on Robust Layout Segmentation in Corporate Documents 11 May 2023 · 0 repositories · arXiv:2305.06553
-
BiRT: Bio-inspired Replay in Vision Transformers for Continual Learning 8 May 2023 · 2 repositories · arXiv:2305.04769Syntology official (archive's flag): 3 ran · 3 ran (of which 2 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 1 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples)
-
Understanding Gaussian Attention Bias of Vision Transformers Using Effective Receptive Fields 8 May 2023 · 1 repository · arXiv:2305.04722
-
Vision Transformer Off-the-Shelf: A Surprising Baseline for Few-Shot Class-Agnostic Counting 8 May 2023 · 1 repository · arXiv:2305.04440
-
Model-Contrastive Federated Domain Adaptation 7 May 2023 · 0 repositories · arXiv:2305.10432
-
FM-ViT: Flexible Modal Vision Transformers for Face Anti-Spoofing 5 May 2023 · 0 repositories · arXiv:2305.03277
-
Reduction of Class Activation Uncertainty with Background Information 5 May 2023 · 2 repositories · arXiv:2305.03238
-
A Vision Transformer Approach for Efficient Near-Field Irregular SAR Super-Resolution 3 May 2023 · 2 repositories · arXiv:2305.02074
-
Glitch in the Matrix: A Large Scale Benchmark for Content Driven Audio-Visual Forgery Detection and Localization 3 May 2023 · 1 repository · arXiv:2305.01979
-
Learngene: Inheriting Condensed Knowledge from the Ancestry Model to Descendant Models 3 May 2023 · 0 repositories · arXiv:2305.02279
-
ARBEx: Attentive Feature Extraction with Reliability Balancing for Robust Facial Expression Learning 2 May 2023 · 1 repository · arXiv:2305.01486
-
AxWin Transformer: A Context-Aware Vision Transformer Backbone with Axial Windows 2 May 2023 · 0 repositories · arXiv:2305.01280
-
Exploring vision transformer layer choosing for semantic segmentation 2 May 2023 · 0 repositories · arXiv:2305.01279
-
Rethinking Boundary Detection in Deep Learning Models for Medical Image Segmentation 1 May 2023 · 1 repository · arXiv:2305.00678
-
An automated end-to-end deep learning-based framework for lung cancer diagnosis by detecting and classifying the lung nodules 28 Apr 2023 · 0 repositories · arXiv:2305.00046
-
DIAMANT: Dual Image-Attention Map Encoders For Medical Image Segmentation 28 Apr 2023 · 0 repositories · arXiv:2304.14571
-
TextDeformer: Geometry Manipulation using Text Guidance 26 Apr 2023 · 1 repository · arXiv:2304.13348Syntology official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 1 where Syntology's instrument failed) · 5 unverified (of 16 harvested samples) · 1 pointer-only (licence)
-
CompletionFormer: Depth Completion with Convolutions and Vision Transformers 25 Apr 2023 · 1 repository · arXiv:2304.13030
-
Augmentation-based Domain Generalization for Semantic Segmentation 24 Apr 2023 · 0 repositories · arXiv:2304.12122
-
Rank Flow Embedding for Unsupervised and Semi-Supervised Manifold Learning 24 Apr 2023 · 1 repository · arXiv:2304.12448
-
Universal Domain Adaptation via Compressive Attention Matching 24 Apr 2023 · 0 repositories · arXiv:2304.11862
-
Vision Transformer for Efficient Chest X-ray and Gastrointestinal Image Classification 23 Apr 2023 · 0 repositories · arXiv:2304.11529
-
Vision Transformers, a new approach for high-resolution and large-scale mapping of canopy heights 22 Apr 2023 · 0 repositories · arXiv:2304.11487
-
DeformableFormer: Classification of Endoscopic Ultrasound Guided Fine Needle Biopsy in Pancreatic Diseases 21 Apr 2023 · 0 repositories · arXiv:2304.10791
-
Contrastive Tuning: A Little Help to Make Masked Autoencoders Forget 20 Apr 2023 · 1 repository · arXiv:2304.10520
-
Text2Seg: Remote Sensing Image Semantic Segmentation via Text-Guided Visual Foundation Models 20 Apr 2023 · 1 repository · arXiv:2304.10597Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples)
-
Boosting Semantic Segmentation with Semantic Boundaries 19 Apr 2023 · 1 repository · arXiv:2304.09427
-
Self-Supervised Learning from Non-Object Centric Images with a Geometric Transformation Sensitive Architecture 17 Apr 2023 · 1 repository · arXiv:2304.08014
-
Synthetic Data from Diffusion Models Improves ImageNet Classification 17 Apr 2023 · 0 repositories · arXiv:2304.08466
-
Transformer with Selective Shuffled Position Embedding and Key-Patch Exchange Strategy for Early Detection of Knee Osteoarthritis 17 Apr 2023 · 0 repositories · arXiv:2304.08364
-
ViPLO: Vision Transformer based Pose-Conditioned Self-Loop Graph for Human-Object Interaction Detection 17 Apr 2023 · 1 repository · arXiv:2304.08114
-
Align-DETR: Enhancing End-to-end Object Detection with Aligned Loss 15 Apr 2023 · 1 repository · arXiv:2304.07527Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
MA-ViT: Modality-Agnostic Vision Transformers for Face Anti-Spoofing 15 Apr 2023 · 0 repositories · arXiv:2304.07549
-
CAD-RADS scoring of coronary CT angiography with Multi-Axis Vision Transformer: a clinically-inspired deep learning pipeline 14 Apr 2023 · 1 repository · arXiv:2304.07277
-
Very high resolution canopy height maps from RGB imagery using self-supervised vision transformer and convolutional decoder trained on Aerial Lidar 14 Apr 2023 · 1 repository · arXiv:2304.07213
-
Uncovering the Inner Workings of STEGO for Safe Unsupervised Semantic Segmentation 14 Apr 2023 · 1 repository · arXiv:2304.07314
-
VISION DIFFMASK: Faithful Interpretation of Vision Transformers with Differentiable Patch Masking 13 Apr 2023 · 1 repository · arXiv:2304.06391
-
RECLIP: Resource-efficient CLIP by Training with Small Images 12 Apr 2023 · 0 repositories · arXiv:2304.06028
-
Towards Evaluating Explanations of Vision Transformers for Medical Imaging 12 Apr 2023 · 1 repository · arXiv:2304.06133
-
A Billion-scale Foundation Model for Remote Sensing Images 11 Apr 2023 · 0 repositories · arXiv:2304.05215
-
MC-ViViT: Multi-branch Classifier-ViViT to detect Mild Cognitive Impairment in older adults using facial videos 11 Apr 2023 · 0 repositories · arXiv:2304.05292
-
Panoramic Image-to-Image Translation 11 Apr 2023 · 0 repositories · arXiv:2304.04960
-
Detection Transformer with Stable Matching 10 Apr 2023 · 2 repositories · arXiv:2304.04742Syntology official (archive's flag): 1 ran · 7 ran (of which 2 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 3 where Syntology's instrument failed) · 2 unverified (of 9 harvested samples) · 1 pointer-only (licence)
-
Slide-Transformer: Hierarchical Vision Transformer with Local Self-Attention 9 Apr 2023 · 1 repository · arXiv:2304.04237Syntology official (archive's flag): 1 ran · 1 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified; the one sample that ran constructed an object rather than computing a result (of 1 harvested sample) · 1 pointer-only (licence)
-
A Cross-Scale Hierarchical Transformer with Correspondence-Augmented Attention for inferring Bird's-Eye-View Semantic Segmentation 7 Apr 2023 · 0 repositories · arXiv:2304.03650
-
From Saliency to DINO: Saliency-guided Vision Transformer for Few-shot Keypoint Detection 6 Apr 2023 · 0 repositories · arXiv:2304.03140
-
InterFormer: Real-time Interactive Image Segmentation 6 Apr 2023 · 1 repository · arXiv:2304.02942
-
R²Former: Unified Retrieval and Reranking Transformer for Place Recognition 6 Apr 2023 · 0 repositories · arXiv:2304.03410
-
Attention Map Guided Transformer Pruning for Edge Device 4 Apr 2023 · 1 repository · arXiv:2304.01452
-
EPVT: Environment-aware Prompt Vision Transformer for Domain Generalization in Skin Lesion Recognition 4 Apr 2023 · 1 repository · arXiv:2304.01508
-
Strong Baselines for Parameter Efficient Few-Shot Fine-tuning 4 Apr 2023 · 0 repositories · arXiv:2304.01917
-
WeakTr: Exploring Plain Vision Transformer for Weakly-supervised Semantic Segmentation 3 Apr 2023 · 1 repository · arXiv:2304.01184Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Rethinking Local Perception in Lightweight Vision Transformer 31 Mar 2023 · 1 repository · arXiv:2303.17803
-
Visual Anomaly Detection via Dual-Attention Transformer and Discriminative Flow 31 Mar 2023 · 1 repository · arXiv:2303.17882
-
If At First You Don't Succeed: Test Time Re-ranking for Zero-shot, Cross-domain Retrieval 30 Mar 2023 · 0 repositories · arXiv:2303.17703
-
MobileInst: Video Instance Segmentation on the Mobile 30 Mar 2023 · 0 repositories · arXiv:2303.17594
-
Streaming Video Model 30 Mar 2023 · 1 repository · arXiv:2303.17228
-
Whether and When does Endoscopy Domain Pretraining Make Sense? 30 Mar 2023 · 1 repository · arXiv:2303.17636
-
Multi-scale Hierarchical Vision Transformer with Cascaded Attention Decoding for Medical Image Segmentation 29 Mar 2023 · 1 repository · arXiv:2303.16892
-
Self-accumulative Vision Transformer for Bone Age Assessment Using the Sauvegrain Method 29 Mar 2023 · 0 repositories · arXiv:2303.16557
-
ASIC: Aligning Sparse in-the-wild Image Collections 28 Mar 2023 · 0 repositories · arXiv:2303.16201