Methods › Computer Vision › Vision Transformers › Vision Transformer › Papers, page 11
Vision Transformer
Papers archive 2025-07-28
archive papers tagged: 2,144 · with a code link: 1,051 · where Syntology ran a sample: 328 (286 with a run with no instrument failure, 42 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (328 of 2,144 tagged: 286 with a run with no instrument failure, 42 where every run was a failure of Syntology's instrument)
Page 11 of 22: papers 1,001 to 1,100 of 2,144, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
IDPL-PFOD2: A New Large-Scale Dataset for Printed Farsi Optical Character Recognition 2 Dec 2023 · 1 repository · arXiv:2312.01177
-
USat: A Unified Self-Supervised Encoder for Multi-Sensor Satellite Imagery 2 Dec 2023 · 1 repository · arXiv:2312.02199
-
BCN: Batch Channel Normalization for Image Classification 1 Dec 2023 · 1 repository · arXiv:2312.00596
-
Deep Unlearning: Fast and Efficient Gradient-free Approach to Class Forgetting 1 Dec 2023 · 1 repository · arXiv:2312.00761Syntology official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 0 where Syntology's instrument failed) · 7 unverified (of 17 harvested samples) · 17 pointer-only (licence)
-
Generative Parameter-Efficient Fine-Tuning 1 Dec 2023 · 1 repository · arXiv:2312.00700Syntology official (archive's flag): 12 ran · 12 ran (of which 0 constructed an object rather than computing a result; 12 with no instrument failure: 0 honoured, 0 violated, 12 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 16 harvested samples) · 10 pointer-only (licence)
-
Improve Supervised Representation Learning with Masked Image Modeling 1 Dec 2023 · 0 repositories · arXiv:2312.00950
-
SynFundus-1M: A High-quality Million-scale Synthetic fundus images Dataset with Fifteen Types of Annotation 1 Dec 2023 · 1 repository · arXiv:2312.00377
-
A Lightweight Clustering Framework for Unsupervised Semantic Segmentation 30 Nov 2023 · 0 repositories · arXiv:2311.18628
-
HiFi Tuner: High-Fidelity Subject-Driven Fine-Tuning for Diffusion Models 30 Nov 2023 · 0 repositories · arXiv:2312.00079
-
Knowledge Transfer from Vision Foundation Models for Efficient Training of Small Task-specific Models 30 Nov 2023 · 1 repository · arXiv:2311.18237
-
Stochastic Vision Transformers with Wasserstein Distance-Aware Attention 30 Nov 2023 · 0 repositories · arXiv:2311.18645
-
Betrayed by Attention: A Simple yet Effective Approach for Self-supervised Video Object Segmentation 29 Nov 2023 · 1 repository · arXiv:2311.17893Syntology official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 2 where Syntology's instrument failed) · 2 unverified (of 12 harvested samples) · 12 pointer-only (licence)
-
Contrastive Vision-Language Alignment Makes Efficient Instruction Learner 29 Nov 2023 · 1 repository · arXiv:2311.17945
-
PViT-6D: Overclocking Vision Transformers for 6D Pose Estimation with Confidence-Level Prediction and Pose Tokens 29 Nov 2023 · 1 repository · arXiv:2311.17504
-
Spherical Frustum Sparse Convolution Network for LiDAR Point Cloud Semantic Segmentation 29 Nov 2023 · 1 repository · arXiv:2311.17491Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 2 violated, 4 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 7 harvested samples) · 2 pointer-only (licence)
-
DyRA: Portable Dynamic Resolution Adjustment Network for Existing Detectors 28 Nov 2023 · 2 repositories · arXiv:2311.17098
-
STR-Cert: Robustness Certification for Deep Text Recognition on Deep Learning Pipelines and Vision Transformers 28 Nov 2023 · 0 repositories · arXiv:2401.05338
-
Machine Learning-Based Jamun Leaf Disease Detection: A Comprehensive Review 27 Nov 2023 · 0 repositories · arXiv:2311.15741
-
ChAda-ViT : Channel Adaptive Attention for Joint Representation Learning of Heterogeneous Microscopy Images 26 Nov 2023 · 2 repositories · arXiv:2311.15264Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 6 harvested samples)
-
Spectro-ViT: A Vision Transformer Model for GABA-edited MRS Reconstruction Using Spectrograms 26 Nov 2023 · 0 repositories · arXiv:2311.15386
-
Ultra-Range Gesture Recognition using a Web-Camera in Human-Robot Interaction 26 Nov 2023 · 0 repositories · arXiv:2311.15361
-
GeoViT: A Versatile Vision Transformer Architecture for Geospatial Image Analysis 24 Nov 2023 · 0 repositories · arXiv:2311.14301
-
TVT: Training-Free Vision Transformer Search on Tiny Datasets 24 Nov 2023 · 0 repositories · arXiv:2311.14337
-
Understanding Self-Supervised Features for Learning Unsupervised Instance Segmentation 24 Nov 2023 · 0 repositories · arXiv:2311.14665
-
FViT-Grasp: Grasping Objects With Using Fast Vision Transformers 23 Nov 2023 · 0 repositories · arXiv:2311.13986
-
HEViTPose: High-Efficiency Vision Transformer for Human Pose Estimation 22 Nov 2023 · 1 repository · arXiv:2311.13615
-
Feature Extraction for Generative Medical Imaging Evaluation: New Evidence Against an Evolving Trend 22 Nov 2023 · 2 repositories · arXiv:2311.13717
-
Unified Classification and Rejection: A One-versus-All Framework 22 Nov 2023 · 1 repository · arXiv:2311.13355Syntology official: harvested, nothing ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 2 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
White-Box Transformers via Sparse Rate Reduction: Compression Is All There Is? 22 Nov 2023 · 1 repository · arXiv:2311.13110Syntology 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
HoVer-UNet: Accelerating HoVerNet with UNet-based multi-class nuclei segmentation via knowledge distillation 21 Nov 2023 · 1 repository · arXiv:2311.12553
-
A Large-Scale Car Parts (LSCP) Dataset for Lightweight Fine-Grained Detection 20 Nov 2023 · 0 repositories · arXiv:2311.11754
-
Disentangling Structure and Appearance in ViT Feature Space 20 Nov 2023 · 0 repositories · arXiv:2311.12193
-
FreeKD: Knowledge Distillation via Semantic Frequency Prompt 20 Nov 2023 · 1 repository · arXiv:2311.12079Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Tiny-VBF: Resource-Efficient Vision Transformer based Lightweight Beamformer for Ultrasound Single-Angle Plane Wave Imaging 20 Nov 2023 · 0 repositories · arXiv:2311.12082
-
Shape-Sensitive Loss for Catheter and Guidewire Segmentation 19 Nov 2023 · 0 repositories · arXiv:2311.11205
-
Semi-supervised ViT knowledge distillation network with style transfer normalization for colorectal liver metastases survival prediction 17 Nov 2023 · 0 repositories · arXiv:2311.10305
-
Improved TokenPose with Sparsity 16 Nov 2023 · 0 repositories · arXiv:2311.09653
-
UnifiedVisionGPT: Streamlining Vision-Oriented AI through Generalized Multimodal Framework 16 Nov 2023 · 1 repository · arXiv:2311.10125
-
ConvNet vs Transformer, Supervised vs CLIP: Beyond ImageNet Accuracy 15 Nov 2023 · 1 repository · arXiv:2311.09215Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
DISTA: Denoising Spiking Transformer with intrinsic plasticity and spatiotemporal attention 15 Nov 2023 · 0 repositories · arXiv:2311.09376
-
Generalizable Imitation Learning Through Pre-Trained Representations 15 Nov 2023 · 0 repositories · arXiv:2311.09350
-
SkelVIT: Consensus of Vision Transformers for a Lightweight Skeleton-Based Action Recognition System 14 Nov 2023 · 0 repositories · arXiv:2311.08094
-
Dual-channel Prototype Network for few-shot Classification of Pathological Images 14 Nov 2023 · 0 repositories · arXiv:2311.07871
-
TSViT: A Time Series Vision Transformer for Fault Diagnosis 12 Nov 2023 · 0 repositories · arXiv:2311.06916
-
Two Stream Scene Understanding on Graph Embedding 12 Nov 2023 · 0 repositories · arXiv:2311.06746
-
Automatic Report Generation for Histopathology images using pre-trained Vision Transformers 10 Nov 2023 · 1 repository · arXiv:2311.06176
-
Glioblastoma Tumor Segmentation using an Ensemble of Vision Transformers 9 Nov 2023 · 1 repository · arXiv:2312.11467
-
Intelligent Cervical Spine Fracture Detection Using Deep Learning Methods 9 Nov 2023 · 0 repositories · arXiv:2311.05708
-
Vision Encoder-Decoder Models for AI Coaching 9 Nov 2023 · 2 repositories · arXiv:2311.16161
-
Beyond Size: How Gradients Shape Pruning Decisions in Large Language Models 8 Nov 2023 · 2 repositories · arXiv:2311.04902Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 1 where Syntology's instrument failed) · 4 unverified (of 7 harvested samples) · 1 pointer-only (licence)
-
FibroVit—Vision transformer-based framework for detection and classification of pulmonary fibrosis from chest CT images 8 Nov 2023 · 1 repository
-
Dual-Stream Attention Transformers for Sewer Defect Classification 7 Nov 2023 · 1 repository · arXiv:2311.16145
-
FusionViT: Hierarchical 3D Object Detection via LiDAR-Camera Vision Transformer Fusion 7 Nov 2023 · 0 repositories · arXiv:2311.03620
-
Lightweight Portrait Matting via Regional Attention and Refinement 7 Nov 2023 · 0 repositories · arXiv:2311.03770
-
A Recent Survey of the Advancements in Deep Learning Techniques for Monkeypox Disease Detection 6 Nov 2023 · 0 repositories · arXiv:2311.10754
-
Asymmetric Masked Distillation for Pre-Training Small Foundation Models 6 Nov 2023 · 0 repositories · arXiv:2311.03149
-
Cal-DETR: Calibrated Detection Transformer 6 Nov 2023 · 1 repository · arXiv:2311.03570Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample)
-
Machine Learning-Based Tea Leaf Disease Detection: A Comprehensive Review 6 Nov 2023 · 0 repositories · arXiv:2311.03240
-
Masking Hyperspectral Imaging Data with Pretrained Models 6 Nov 2023 · 1 repository · arXiv:2311.03053
-
SugarViT -- Multi-objective Regression of UAV Images with Vision Transformers and Deep Label Distribution Learning Demonstrated on Disease Severity Prediction in Sugar Beet 6 Nov 2023 · 0 repositories · arXiv:2311.03076
-
Rotation Invariant Transformer for Recognizing Object in UAVs 5 Nov 2023 · 3 repositories · arXiv:2311.02559
-
ProS: Facial Omni-Representation Learning via Prototype-based Self-Distillation 3 Nov 2023 · 0 repositories · arXiv:2311.01929
-
Distilling Knowledge from CNN-Transformer Models for Enhanced Human Action Recognition 2 Nov 2023 · 0 repositories · arXiv:2311.01283
-
Efficient Vision Transformer for Accurate Traffic Sign Detection 2 Nov 2023 · 0 repositories · arXiv:2311.01429
-
Scattering Vision Transformer: Spectral Mixing Matters 2 Nov 2023 · 0 repositories · arXiv:2311.01310
-
Patch-Based Deep Unsupervised Image Segmentation using Graph Cuts 1 Nov 2023 · 0 repositories · arXiv:2311.01475
-
Addressing Limitations of State-Aware Imitation Learning for Autonomous Driving 31 Oct 2023 · 0 repositories · arXiv:2310.20650
-
In Search of Lost Online Test-time Adaptation: A Survey 31 Oct 2023 · 1 repository · arXiv:2310.20199Syntology official (archive's flag): 14 ran · 14 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 5 where Syntology's instrument failed) · 3 unverified (of 17 harvested samples) · 7 pointer-only (licence)
-
AViTMP: A Tracking-Specific Transformer for Single-Branch Visual Tracking 30 Oct 2023 · 1 repository · arXiv:2310.19542
-
MIST: Medical Image Segmentation Transformer with Convolutional Attention Mixing (CAM) Decoder 30 Oct 2023 · 1 repository · arXiv:2310.19898
-
Promise:Prompt-driven 3D Medical Image Segmentation Using Pretrained Image Foundation Models 30 Oct 2023 · 1 repository · arXiv:2310.19721
-
Uncovering Prototypical Knowledge for Weakly Open-Vocabulary Semantic Segmentation 29 Oct 2023 · 1 repository · arXiv:2310.19001Syntology 5 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
Patch-Wise Self-Supervised Visual Representation Learning: A Fine-Grained Approach 28 Oct 2023 · 1 repository · arXiv:2310.18651
-
A Self-Supervised Approach to Land Cover Segmentation 27 Oct 2023 · 0 repositories · arXiv:2310.18251
-
Qilin-Med-VL: Towards Chinese Large Vision-Language Model for General Healthcare 27 Oct 2023 · 1 repository · arXiv:2310.17956Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
A No-Reference Quality Assessment Method for Digital Human Head 25 Oct 2023 · 0 repositories · arXiv:2310.16732
-
LLM-FP4: 4-Bit Floating-Point Quantized Transformers 25 Oct 2023 · 1 repository · arXiv:2310.16836Syntology official (archive's flag): 3 ran · 3 ran (of which 3 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified; every one of the 3 samples that ran constructed an object rather than computing a result (of 3 harvested samples)
-
SAMCLR: Contrastive pre-training on complex scenes using SAM for view sampling 23 Oct 2023 · 0 repositories · arXiv:2310.14736
-
CXR-LLAVA: a multimodal large language model for interpreting chest X-ray images 22 Oct 2023 · 1 repository · arXiv:2310.18341Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Concept-based Anomaly Detection in Retail Stores for Automatic Correction using Mobile Robots 21 Oct 2023 · 0 repositories · arXiv:2310.14063
-
Exploring Driving Behavior for Autonomous Vehicles Based on Gramian Angular Field Vision Transformer 21 Oct 2023 · 0 repositories · arXiv:2310.13906
-
Heart Disease Detection using Vision-Based Transformer Models from ECG Images 19 Oct 2023 · 0 repositories · arXiv:2310.12630
-
LeTFuser: Light-weight End-to-end Transformer-Based Sensor Fusion for Autonomous Driving with Multi-Task Learning 19 Oct 2023 · 1 repository · arXiv:2310.13135
-
Minimalist and High-Performance Semantic Segmentation with Plain Vision Transformers 19 Oct 2023 · 1 repository · arXiv:2310.12755
-
FixPix: Fixing Bad Pixels using Deep Learning 18 Oct 2023 · 0 repositories · arXiv:2310.11637
-
Tailoring Adversarial Attacks on Deep Neural Networks for Targeted Class Manipulation Using DeepFool Algorithm 18 Oct 2023 · 0 repositories · arXiv:2310.13019
-
Image Augmentation with Controlled Diffusion for Weakly-Supervised Semantic Segmentation 15 Oct 2023 · 0 repositories · arXiv:2310.09760
-
MoEmo Vision Transformer: Integrating Cross-Attention and Movement Vectors in 3D Pose Estimation for HRI Emotion Detection 15 Oct 2023 · 1 repository · arXiv:2310.09757
-
Top-K Pooling with Patch Contrastive Learning for Weakly-Supervised Semantic Segmentation 15 Oct 2023 · 0 repositories · arXiv:2310.09828
-
Vision Transformers increase efficiency of 3D cardiac CT multi-label segmentation 13 Oct 2023 · 1 repository · arXiv:2310.09099
-
From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models 13 Oct 2023 · 1 repository · arXiv:2310.08825
-
PaLI-3 Vision Language Models: Smaller, Faster, Stronger 13 Oct 2023 · 1 repository · arXiv:2310.09199Syntology 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 3 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 1 pointer-only (licence)
-
Distilling Efficient Vision Transformers from CNNs for Semantic Segmentation 11 Oct 2023 · 0 repositories · arXiv:2310.07265
-
Does resistance to style-transfer equal Global Shape Bias? Measuring network sensitivity to global shape configuration 11 Oct 2023 · 0 repositories · arXiv:2310.07555
-
ProtoHPE: Prototype-guided High-frequency Patch Enhancement for Visible-Infrared Person Re-identification 11 Oct 2023 · 0 repositories · arXiv:2310.07552
-
PtychoDV: Vision Transformer-Based Deep Unrolling Network for Ptychographic Image Reconstruction 11 Oct 2023 · 1 repository · arXiv:2310.07504
-
Computational Pathology at Health System Scale -- Self-Supervised Foundation Models from Three Billion Images 10 Oct 2023 · 0 repositories · arXiv:2310.07033
-
Efficient Adaptation of Large Vision Transformer via Adapter Re-Composing 10 Oct 2023 · 1 repository · arXiv:2310.06234Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 7 harvested samples) · 4 pointer-only (licence)
-
EViT: An Eagle Vision Transformer with Bi-Fovea Self-Attention 10 Oct 2023 · 1 repository · arXiv:2310.06629
-
Learning Stackable and Skippable LEGO Bricks for Efficient, Reconfigurable, and Variable-Resolution Diffusion Modeling 10 Oct 2023 · 1 repository · arXiv:2310.06389Syntology official (archive's flag): 13 ran · 13 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 3 honoured, 0 violated, 5 with no contract checked; 5 where Syntology's instrument failed) · 4 unverified (of 17 harvested samples) · 8 pointer-only (licence)