Methods › Computer Vision › Vision Transformers › Vision Transformer › Papers, page 2
Vision Transformer
Papers archive 2025-07-28
archive papers tagged: 2,144 · with a code link: 1,051 · where Syntology ran a sample: 328 (286 with a run with no instrument failure, 42 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (328 of 2,144 tagged: 286 with a run with no instrument failure, 42 where every run was a failure of Syntology's instrument)
Page 2 of 22: papers 101 to 200 of 2,144, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Advanced Chest X-Ray Analysis via Transformer-Based Image Descriptors and Cross-Model Attention Mechanism 23 Apr 2025 · 0 repositories · arXiv:2504.16774
-
Quantum Doubly Stochastic Transformers 22 Apr 2025 · 0 repositories · arXiv:2504.16275Syntology 3 ran (of which 2 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples) · 1 pointer-only (licence)
-
Automated Measurement of Eczema Severity with Self-Supervised Learning 21 Apr 2025 · 0 repositories · arXiv:2504.15193
-
Distribution-aware Dataset Distillation for Efficient Image Restoration 21 Apr 2025 · 0 repositories · arXiv:2504.14826
-
ECViT: Efficient Convolutional Vision Transformer with Local-Attention and Multi-scale Stages 21 Apr 2025 · 1 repository · arXiv:2504.14825
-
Impact of Latent Space Dimension on IoT Botnet Detection Performance: VAE-Encoder Versus ViT-Encoder 21 Apr 2025 · 0 repositories · arXiv:2504.14879
-
MSAD-Net: Multiscale and Spatial Attention-based Dense Network for Lung Cancer Classification 20 Apr 2025 · 0 repositories · arXiv:2504.14626
-
6G WavesFM: A Foundation Model for Sensing, Communication, and Localization 18 Apr 2025 · 0 repositories · arXiv:2504.14100
-
A Deep Learning-Based Supervised Transfer Learning Framework for DOA Estimation with Array Imperfections 18 Apr 2025 · 1 repository · arXiv:2504.13394
-
Collective Learning Mechanism based Optimal Transport Generative Adversarial Network for Non-parallel Voice Conversion 18 Apr 2025 · 0 repositories · arXiv:2504.13791
-
CytoFM: The first cytology foundation model 18 Apr 2025 · 0 repositories · arXiv:2504.13402
-
DenSe-AdViT: A novel Vision Transformer for Dense SAR Object Detection 18 Apr 2025 · 0 repositories · arXiv:2504.13638
-
LoRA-Based Continual Learning with Constraints on Critical Parameter Changes 18 Apr 2025 · 1 repository · arXiv:2504.13407
-
Putting the Segment Anything Model to the Test with 3D Knee MRI - A Comparison with State-of-the-Art Performance 17 Apr 2025 · 1 repository · arXiv:2504.13340
-
Human Aligned Compression for Robust Models 16 Apr 2025 · 1 repository · arXiv:2504.12255
-
Zooming In on Fakes: A Novel Dataset for Localized AI-Generated Image Detection with Forgery Amplification Approach 16 Apr 2025 · 1 repository · arXiv:2504.11922
-
A Decade of Wheat Mapping for Lebanon 15 Apr 2025 · 0 repositories · arXiv:2504.11366
-
AFiRe: Anatomy-Driven Self-Supervised Learning for Fine-Grained Representation in Radiographic Images 15 Apr 2025 · 1 repository · arXiv:2504.10972
-
Embedding Radiomics into Vision Transformers for Multimodal Medical Image Classification 15 Apr 2025 · 0 repositories · arXiv:2504.10916
-
Differentially Private 2D Human Pose Estimation 14 Apr 2025 · 0 repositories · arXiv:2504.10190
-
Integrating Vision and Location with Transformers: A Multimodal Deep Learning Framework for Medical Wound Analysis 14 Apr 2025 · 0 repositories · arXiv:2504.10452
-
Self-Controlled Dynamic Expansion Model for Continual Learning 14 Apr 2025 · 0 repositories · arXiv:2504.10561
-
The Scalability of Simplicity: Empirical Analysis of Vision-Language Learning with a Single Transformer 14 Apr 2025 · 1 repository · arXiv:2504.10462Syntology official (archive's flag): 2 ran · 2 ran (of which 2 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 7 unverified; every one of the 2 samples that ran constructed an object rather than computing a result (of 9 harvested samples)
-
Learning Occlusion-Robust Vision Transformers for Real-Time UAV Tracking 12 Apr 2025 · 1 repository · arXiv:2504.09228Syntology official (archive's flag): 1 ran · 1 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified; the one sample that ran constructed an object rather than computing a result (of 1 harvested sample)
-
Multi-scale Activation, Refinement, and Aggregation: Exploring Diverse Cues for Fine-Grained Bird Recognition 12 Apr 2025 · 0 repositories · arXiv:2504.09215
-
Adaptive Additive Parameter Updates of Vision Transformers for Few-Shot Continual Learning 11 Apr 2025 · 0 repositories · arXiv:2504.08982
-
Hypergraph Vision Transformers: Images are More than Nodes, More than Edges 11 Apr 2025 · 0 repositories · arXiv:2504.08710
-
SARFormer -- An Acquisition Parameter Aware Vision Transformer for Synthetic Aperture Radar Data 11 Apr 2025 · 0 repositories · arXiv:2504.08441
-
Steering CLIP's vision transformer with sparse autoencoders 11 Apr 2025 · 0 repositories · arXiv:2504.08729
-
Breaking the Barriers: Video Vision Transformers for Word-Level Sign Language Recognition 10 Apr 2025 · 0 repositories · arXiv:2504.07792
-
Deep Learning Meets Teleconnections: Improving S2S Predictions for European Winter Weather 10 Apr 2025 · 1 repository · arXiv:2504.07625
-
Heart Failure Prediction using Modal Decomposition and Masked Autoencoders for Scarce Echocardiography Databases 10 Apr 2025 · 1 repository · arXiv:2504.07606
-
Novel Pooling-based VGG-Lite for Pneumonia and Covid-19 Detection from Imbalanced Chest X-Ray Datasets 10 Apr 2025 · 0 repositories · arXiv:2504.07468
-
Gaze-Guided Learning: Avoiding Shortcut Bias in Visual Classification 8 Apr 2025 · 1 repository · arXiv:2504.05583
-
HRMedSeg: Unlocking High-resolution Medical Image segmentation via Memory-efficient Attention Modeling 8 Apr 2025 · 1 repository · arXiv:2504.06205
-
PromptHMR: Promptable Human Mesh Recovery 8 Apr 2025 · 1 repository · arXiv:2504.06397
-
EMF: Event Meta Formers for Event-based Real-time Traffic Object Detection 5 Apr 2025 · 0 repositories · arXiv:2504.04124
-
Resilience of Vision Transformers for Domain Generalisation in the Presence of Out-of-Distribution Noisy Images 5 Apr 2025 · 0 repositories · arXiv:2504.04225
-
AdaViT: Adaptive Vision Transformer for Flexible Pretrain and Finetune with Variable 3D Medical Image Modalities 4 Apr 2025 · 0 repositories · arXiv:2504.03589
-
AC-LoRA: Auto Component LoRA for Personalized Artistic Style Image Generation 3 Apr 2025 · 0 repositories · arXiv:2504.02231
-
Beyond Conventional Transformers: The Medical X-ray Attention (MXA) Block for Improved Multi-Label Diagnosis Using Knowledge Distillation 3 Apr 2025 · 1 repository · arXiv:2504.02277
-
F-ViTA: Foundation Model Guided Visible to Thermal Translation 3 Apr 2025 · 1 repository · arXiv:2504.02801
-
GPTAQ: Efficient Finetuning-Free Quantization for Asymmetric Calibration 3 Apr 2025 · 2 repositories · arXiv:2504.02692Syntology official: harvested, nothing ran · 0 ran · 2 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
HGFormer: Topology-Aware Vision Transformer with HyperGraph Learning 3 Apr 2025 · 0 repositories · arXiv:2504.02440
-
HQViT: Hybrid Quantum Vision Transformer for Image Classification 3 Apr 2025 · 0 repositories · arXiv:2504.02730
-
Semiconductor Wafer Map Defect Classification with Tiny Vision Transformers 3 Apr 2025 · 0 repositories · arXiv:2504.02494
-
Prompt-Guided Attention Head Selection for Focus-Oriented Image Retrieval 2 Apr 2025 · 0 repositories · arXiv:2504.01348
-
UniViTAR: Unified Vision Transformer with Native Resolution 2 Apr 2025 · 0 repositories · arXiv:2504.01792
-
CellVTA: Enhancing Vision Foundation Models for Accurate Cell Segmentation and Classification 1 Apr 2025 · 1 repository · arXiv:2504.00784
-
MergeVQ: A Unified Framework for Visual Generation and Representation with Disentangled Token Merging and Quantization 1 Apr 2025 · 1 repository · arXiv:2504.00999
-
QSViT: A Methodology for Quantizing Spiking Vision Transformers 1 Apr 2025 · 0 repositories · arXiv:2504.00948
-
Conformal uncertainty quantification to evaluate predictive fairness of foundation AI model for skin lesion classes across patient demographics 31 Mar 2025 · 0 repositories · arXiv:2503.23819
-
Foundation Models For Seismic Data Processing: An Extensive Review 31 Mar 2025 · 1 repository · arXiv:2503.24166
-
A Lightweight Image Super-Resolution Transformer Trained on Low-Resolution Images Only 30 Mar 2025 · 1 repository · arXiv:2503.23265
-
Efficient Adaptation For Remote Sensing Visual Grounding 29 Mar 2025 · 0 repositories · arXiv:2503.23083
-
Large Self-Supervised Models Bridge the Gap in Domain Adaptive Object Detection 29 Mar 2025 · 1 repository · arXiv:2503.23220
-
Z-SASLM: Zero-Shot Style-Aligned SLI Blending Latent Manipulation 29 Mar 2025 · 1 repository · arXiv:2503.23234
-
ReCoM: Realistic Co-Speech Motion Generation with Recurrent Embedded Transformer 27 Mar 2025 · 0 repositories · arXiv:2503.21847
-
Face Spoofing Detection using Deep Learning 25 Mar 2025 · 1 repository · arXiv:2503.19223
-
RoboFlamingo-Plus: Fusion of Depth and RGB Perception with Vision-Language Models for Enhanced Robotic Manipulation 25 Mar 2025 · 0 repositories · arXiv:2503.19510
-
Surg-3M: A Dataset and Foundation Model for Perception in Surgical Settings 25 Mar 2025 · 1 repository · arXiv:2503.19740
-
Chirp Localization via Fine-Tuned Transformer Model: A Proof-of-Concept Study 24 Mar 2025 · 0 repositories · arXiv:2503.22713
-
Image-to-Text for Medical Reports Using Adaptive Co-Attention and Triple-LSTM Module 24 Mar 2025 · 0 repositories · arXiv:2503.18297
-
SPMTrack: Spatio-Temporal Parameter-Efficient Fine-Tuning with Mixture of Experts for Scalable Visual Tracking 24 Mar 2025 · 1 repository · arXiv:2503.18338Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
PathoHR: Breast Cancer Survival Prediction on High-Resolution Pathological Images 23 Mar 2025 · 1 repository · arXiv:2503.17970
-
Automated diagnosis of lung diseases using vision transformer: a comparative study on chest x-ray classification 22 Mar 2025 · 0 repositories · arXiv:2503.18973
-
EMPLACE: Self-Supervised Urban Scene Change Detection 22 Mar 2025 · 1 repository · arXiv:2503.17716Syntology official (archive's flag): 4 ran · 4 ran (of which 4 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified; every one of the 4 samples that ran constructed an object rather than computing a result (of 7 harvested samples) · 7 pointer-only (licence)
-
Serial Low-rank Adaptation of Vision Transformer 22 Mar 2025 · 0 repositories · arXiv:2503.17750
-
Feature-Based Dual Visual Feature Extraction Model for Compound Multimodal Emotion Recognition 21 Mar 2025 · 1 repository · arXiv:2503.17453
-
Vision Transformer Based Semantic Communications for Next Generation Wireless Networks 21 Mar 2025 · 0 repositories · arXiv:2503.17275
-
Hyperspectral Imaging for Identifying Foreign Objects on Pork Belly 20 Mar 2025 · 0 repositories · arXiv:2503.16086
-
A-SCoRe: Attention-based Scene Coordinate Regression for wide-ranging scenarios 18 Mar 2025 · 1 repository · arXiv:2503.13982
-
Dynamic Accumulated Attention Map for Interpreting Evolution of Decision-Making in Vision Transformer 18 Mar 2025 · 1 repository · arXiv:2503.14640
-
Text-Guided Image Invariant Feature Learning for Robust Image Watermarking 18 Mar 2025 · 0 repositories · arXiv:2503.13805
-
Advancing Chronic Tuberculosis Diagnostics Using Vision-Language Models: A Multi modal Framework for Precision Analysis 17 Mar 2025 · 0 repositories · arXiv:2503.14536
-
An interpretable approach to automating the assessment of biofouling in video footage 17 Mar 2025 · 1 repository · arXiv:2503.12875
-
Towards Scalable Foundation Model for Multi-modal and Hyperspectral Geospatial Data 17 Mar 2025 · 0 repositories · arXiv:2503.12843
-
Fourier-Based 3D Multistage Transformer for Aberration Correction in Multicellular Specimens 16 Mar 2025 · 2 repositories · arXiv:2503.12593
-
Semantic Matters: Multimodal Features for Affective Analysis 16 Mar 2025 · 0 repositories · arXiv:2504.11460
-
Asynchronous Sharpness-Aware Minimization For Fast and Accurate Deep Learning 14 Mar 2025 · 0 repositories · arXiv:2503.11147
-
BEVDiffLoc: End-to-End LiDAR Global Localization in BEV View based on Diffusion Model 14 Mar 2025 · 1 repository · arXiv:2503.11372
-
DynRsl-VLM: Enhancing Autonomous Driving Perception with Dynamic Resolution Vision-Language Models 14 Mar 2025 · 0 repositories · arXiv:2503.11265
-
CountPath: Automating Fragment Counting in Digital Pathology 13 Mar 2025 · 0 repositories · arXiv:2503.10520
-
Multi-Domain Biometric Recognition using Body Embeddings 13 Mar 2025 · 0 repositories · arXiv:2503.10931
-
Robustness Tokens: Towards Adversarial Robustness of Transformers 13 Mar 2025 · 1 repository · arXiv:2503.10191
-
CleverDistiller: Simple and Spatially Consistent Cross-modal Distillation 12 Mar 2025 · 0 repositories · arXiv:2503.09878
-
Mapping fMRI Signal and Image Stimuli in an Artificial Neural Network Latent Space: Bringing Artificial and Natural Minds Together 12 Mar 2025 · 0 repositories · arXiv:2503.19923
-
Object-Aware DINO (Oh-A-Dino): Enhancing Self-Supervised Representations for Multi-Object Instance Retrieval 12 Mar 2025 · 0 repositories · arXiv:2503.09867
-
KAN-Mixers: a new deep learning architecture for image classification 11 Mar 2025 · 0 repositories · arXiv:2503.08939
-
TransECG: Leveraging Transformers for Explainable ECG Re-identification Risk Analysis 11 Mar 2025 · 0 repositories · arXiv:2503.13495
-
Vision Transformer for Intracranial Hemorrhage Classification in CT Scans Using an Entropy-Aware Fuzzy Integral Strategy for Adaptive Scan-Level Decision Fusion 11 Mar 2025 · 0 repositories · arXiv:2503.08609
-
A Data-Centric Revisit of Pre-Trained Vision Models for Robot Learning 10 Mar 2025 · 1 repository · arXiv:2503.06960Syntology official (archive's flag): 3 ran · 3 ran (of which 2 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples)
-
Visual and Text Prompt Segmentation: A Novel Multi-Model Framework for Remote Sensing 10 Mar 2025 · 0 repositories · arXiv:2503.07911
-
GroMo: Plant Growth Modeling with Multiview Images 9 Mar 2025 · 1 repository · arXiv:2503.06608
-
Seeing Delta Parameters as JPEG Images: Data-Free Delta Compression with Discrete Cosine Transform 9 Mar 2025 · 0 repositories · arXiv:2503.06676
-
Lightweight Software Kernels and Hardware Extensions for Efficient Sparse Deep Neural Networks on Microcontrollers 8 Mar 2025 · 0 repositories · arXiv:2503.06183
-
GBT-SAM: Adapting a Foundational Deep Learning Model for Generalizable Brain Tumor Segmentation via Efficient Integration of Multi-Parametric MRI Data 6 Mar 2025 · 1 repository · arXiv:2503.04325
-
Toward Lightweight and Fast Decoders for Diffusion Models in Image and Video Generation 6 Mar 2025 · 1 repository · arXiv:2503.04871
-
AHCPTQ: Accurate and Hardware-Compatible Post-Training Quantization for Segment Anything Model 5 Mar 2025 · 0 repositories · arXiv:2503.03088
-
DTU-Net: A Multi-Scale Dilated Transformer Network for Nonlinear Hyperspectral Unmixing 5 Mar 2025 · 0 repositories · arXiv:2503.03465