Methods › Computer Vision › Vision Transformers › Vision Transformer › Papers, page 4
Vision Transformer
Papers archive 2025-07-28
archive papers tagged: 2,144 · with a code link: 1,051 · where Syntology ran a sample: 328 (286 with a run with no instrument failure, 42 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (328 of 2,144 tagged: 286 with a run with no instrument failure, 42 where every run was a failure of Syntology's instrument)
Page 4 of 22: papers 301 to 400 of 2,144, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Efficient and Accurate Tuberculosis Diagnosis: Attention Residual U-Net and Vision Transformer Based Detection Framework 7 Jan 2025 · 0 repositories · arXiv:2501.03538
-
DeTrack: In-model Latent Denoising Learning for Visual Object Tracking 5 Jan 2025 · 0 repositories · arXiv:2501.02467
-
Towards Hard and Soft Shadow Removal via Dual-Branch Separation Network and Vision Transformer 3 Jan 2025 · 0 repositories · arXiv:2501.01864
-
Adaptive Part Learning for Fine-Grained Generalized Category Discovery: A Plug-and-Play Enhancement 1 Jan 2025 · 0 repositories
-
BOE-ViT: Boosting Orientation Estimation with Equivariance in Self-Supervised 3D Subtomogram Alignment 1 Jan 2025 · 0 repositories
-
Closest Neighbors are Harmful for Lightweight Masked Auto-encoders 1 Jan 2025 · 1 repository
-
Correlative and Discriminative Label Grouping for Multi-Label Visual Prompt Tuning 1 Jan 2025 · 0 repositories
-
IceDiff: High Resolution and High-Quality Arctic Sea Ice Forecasting with Generative Diffusion Prior 1 Jan 2025 · 0 repositories
-
Less Attention is More: Prompt Transformer for Generalized Category Discovery 1 Jan 2025 · 0 repositories
-
Prompt-CAM: Making Vision Transformers Interpretable for Fine-Grained Analysis 1 Jan 2025 · 1 repository
-
RADIOv2.5: Improved Baselines for Agglomerative Vision Foundation Models 1 Jan 2025 · 0 repositories
-
Self-Supervised Cross-View Correspondence with Predictive Cycle Consistency 1 Jan 2025 · 0 repositories
-
A Study on Context Length and Efficient Transformers for Biomedical Image Analysis 31 Dec 2024 · 0 repositories · arXiv:2501.00619
-
Advanced Lung Nodule Segmentation and Classification for Early Detection of Lung Cancer using SAM and Transfer Learning 31 Dec 2024 · 0 repositories · arXiv:2501.00586
-
MATEY: multiscale adaptive foundation models for spatiotemporal physical systems 29 Dec 2024 · 0 repositories · arXiv:2412.20601
-
Distilled Transformers with Locally Enhanced Global Representations for Face Forgery Detection 28 Dec 2024 · 0 repositories · arXiv:2412.20156
-
SegKAN: High-Resolution Medical Image Segmentation with Long-Distance Dependencies 28 Dec 2024 · 1 repository · arXiv:2412.19990
-
VisTabNet: Adapting Vision Transformers for Tabular Data 28 Dec 2024 · 1 repository · arXiv:2501.00057
-
Dual Channel Multi-Attention in ViT for Biometric Authentication using Forehead Subcutaneous Vein Pattern and Periocular Pattern 26 Dec 2024 · 0 repositories · arXiv:2412.19160
-
MTCAE-DFER: Multi-Task Cascaded Autoencoder for Dynamic Facial Expression Recognition 25 Dec 2024 · 1 repository · arXiv:2412.18988
-
Unified Local and Global Attention Interaction Modeling for Vision Transformers 25 Dec 2024 · 0 repositories · arXiv:2412.18778
-
AutoSculpt: A Pattern-based Model Auto-pruning Framework Using Reinforcement Learning and Graph Learning 24 Dec 2024 · 0 repositories · arXiv:2412.18091
-
ERVD: An Efficient and Robust ViT-Based Distillation Framework for Remote Sensing Image Retrieval 24 Dec 2024 · 1 repository · arXiv:2412.18136
-
Comprehensive Multi-Modal Prototypes are Simple and Effective Classifiers for Vast-Vocabulary Object Detection 23 Dec 2024 · 1 repository · arXiv:2412.17800
-
Edge-AI for Agriculture: Lightweight Vision Models for Disease Detection in Resource-Limited Settings 23 Dec 2024 · 0 repositories · arXiv:2412.18635
-
Enhancing Contrastive Learning Inspired by the Philosophy of "The Blind Men and the Elephant" 21 Dec 2024 · 1 repository · arXiv:2412.16522
-
Object Detection Approaches to Identifying Hand Images with High Forensic Values 21 Dec 2024 · 0 repositories · arXiv:2412.16431
-
Sensitive Image Classification by Vision Transformers 21 Dec 2024 · 0 repositories · arXiv:2412.16446
-
SeagrassFinder: Deep Learning for Eelgrass Detection and Coverage Estimation in the Wild 20 Dec 2024 · 0 repositories · arXiv:2412.16147
-
Adaptive Prompt Tuning: Vision Guided Prompt Tuning with Cross-Attention for Fine-Grained Few-Shot Learning 19 Dec 2024 · 0 repositories · arXiv:2412.14640
-
Can We Get Rid of Handcrafted Feature Extractors? SparseViT: Nonsemantics-Centered, Parameter-Efficient Image Manipulation Localization through Spare-Coding Transformer 19 Dec 2024 · 1 repository · arXiv:2412.14598Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
DiffSim: Taming Diffusion Models for Evaluating Visual Similarity 19 Dec 2024 · 1 repository · arXiv:2412.14580
-
Jet: A Modern Transformer-Based Normalizing Flow 19 Dec 2024 · 0 repositories · arXiv:2412.15129
-
A LoRA is Worth a Thousand Pictures 16 Dec 2024 · 0 repositories · arXiv:2412.12048
-
MoRe: Class Patch Attention Needs Regularization for Weakly Supervised Semantic Segmentation 15 Dec 2024 · 1 repository · arXiv:2412.11076
-
One-Shot Multilingual Font Generation Via ViT 15 Dec 2024 · 0 repositories · arXiv:2412.11342
-
ManipGPT: Is Affordance Segmentation by Large Vision Models Enough for Articulated Object Manipulation? 13 Dec 2024 · 0 repositories · arXiv:2412.10050
-
T-GMSI: A transformer-based generative model for spatial interpolation under sparse measurements 13 Dec 2024 · 0 repositories · arXiv:2412.09886
-
VibrantVS: A high-resolution multi-task transformer for forest canopy height estimation 13 Dec 2024 · 0 repositories · arXiv:2412.10351
-
A Novel Ensemble-Based Deep Learning Model with Explainable AI for Accurate Kidney Disease Diagnosis 12 Dec 2024 · 0 repositories · arXiv:2412.09472
-
Advancing Attribution-Based Neural Network Explainability through Relative Absolute Magnitude Layer-Wise Relevance Propagation and Multi-Component Evaluation 12 Dec 2024 · 1 repository · arXiv:2412.09311
-
From Noise to Nuance: Advances in Deep Generative Image Models 12 Dec 2024 · 0 repositories · arXiv:2412.09656
-
Selective Visual Prompting in Vision Mamba 12 Dec 2024 · 1 repository · arXiv:2412.08947Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Sensing for Space Safety and Sustainability: A Deep Learning Approach with Vision Transformers 12 Dec 2024 · 0 repositories · arXiv:2412.08913
-
Vision Transformers for Efficient Indoor Pathloss Radio Map Prediction 12 Dec 2024 · 0 repositories · arXiv:2412.09507
-
SAM-Mamba: Mamba Guided SAM Architecture for Generalized Zero-Shot Polyp Segmentation 11 Dec 2024 · 1 repository · arXiv:2412.08482
-
RADIO Amplified: Improved Baselines for Agglomerative Vision Foundation Models 10 Dec 2024 · 1 repository · arXiv:2412.07679
-
Bridging the Divide: Reconsidering Softmax and Linear Attention 9 Dec 2024 · 1 repository · arXiv:2412.06590Syntology official (archive's flag): 17 ran · 17 ran (of which 0 constructed an object rather than computing a result; 13 with no instrument failure: 1 honoured, 0 violated, 12 with no contract checked; 4 where Syntology's instrument failed) · 5 unverified (of 22 harvested samples) · 22 pointer-only (licence)
-
Inverting Transformer-based Vision Models 9 Dec 2024 · 2 repositories · arXiv:2412.06534
-
Knowledge Transfer and Domain Adaptation for Fine-Grained Remote Sensing Image Segmentation 9 Dec 2024 · 1 repository · arXiv:2412.06664
-
Open-Vocabulary High-Resolution 3D (OVHR3D) Data Segmentation and Annotation Framework 9 Dec 2024 · 0 repositories · arXiv:2412.06268
-
ZeroKey: Point-Level Reasoning and Zero-Shot 3D Keypoint Detection from Large Language Models 9 Dec 2024 · 0 repositories · arXiv:2412.06292
-
Enhancing Content Representation for AR Image Quality Assessment Using Knowledge Distillation 8 Dec 2024 · 0 repositories · arXiv:2412.06003
-
Paddy Disease Detection and Classification Using Computer Vision Techniques: A Mobile Application to Detect Paddy Disease 8 Dec 2024 · 0 repositories · arXiv:2412.05996
-
Vision Transformer-based Semantic Communications With Importance-Aware Quantization 8 Dec 2024 · 0 repositories · arXiv:2412.06038
-
RefSAM3D: Adapting SAM with Cross-modal Reference for 3D Medical Image Segmentation 7 Dec 2024 · 0 repositories · arXiv:2412.05605
-
PCTreeS: 3D Point Cloud Tree Species Classification Using Airborne LiDAR Images 6 Dec 2024 · 0 repositories · arXiv:2412.04714
-
Superpixel Tokenization for Vision Transformers: Preserving Semantic Integrity in Visual Tokens 6 Dec 2024 · 1 repository · arXiv:2412.04680Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Automated LaTeX Code Generation from Handwritten Math Expressions Using Vision Transformer 5 Dec 2024 · 0 repositories · arXiv:2412.03853
-
Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion 5 Dec 2024 · 1 repository · arXiv:2412.04424
-
DIVE: Taming DINO for Subject-Driven Video Editing 4 Dec 2024 · 0 repositories · arXiv:2412.03347
-
GraPix: Exploring Graph Modularity Optimization for Unsupervised Pixel Clustering 4 Dec 2024 · 1 repository
-
Multi-Branch Mutual-Distillation Transformer for EEG-Based Seizure Subtype Classification 4 Dec 2024 · 0 repositories · arXiv:2412.15224
-
FCL-ViT: Task-Aware Attention Tuning for Continual Learning 3 Dec 2024 · 0 repositories · arXiv:2412.02509
-
Global Average Feature Augmentation for Robust Semantic Segmentation with Transformers 2 Dec 2024 · 0 repositories · arXiv:2412.01941
-
Mutli-View 3D Reconstruction using Knowledge Distillation 2 Dec 2024 · 1 repository · arXiv:2412.02039
-
Categorical Keypoint Positional Embedding for Robust Animal Re-Identification 1 Dec 2024 · 0 repositories · arXiv:2412.00818
-
Visual Modality Prompt for Adapting Vision-Language Object Detectors 1 Dec 2024 · 1 repository · arXiv:2412.00622
-
Automatic Prompt Generation and Grounding Object Detection for Zero-Shot Image Anomaly Detection 28 Nov 2024 · 0 repositories · arXiv:2411.19220
-
CLIP meets DINO for Tuning Zero-Shot Classifier using Unlabeled Image Collections 28 Nov 2024 · 1 repository · arXiv:2411.19346
-
Efficient Track Anything 28 Nov 2024 · 1 repository · arXiv:2411.18933Syntology 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Enhancing Parameter-Efficient Fine-Tuning of Vision Transformers through Frequency-Based Adaptation 28 Nov 2024 · 1 repository · arXiv:2411.19297
-
MaskRIS: Semantic Distortion-aware Data Augmentation for Referring Image Segmentation 28 Nov 2024 · 1 repository · arXiv:2411.19067
-
Talking to DINO: Bridging Self-Supervised Vision Backbones with Language for Open-Vocabulary Segmentation 28 Nov 2024 · 1 repository · arXiv:2411.19331Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 10 harvested samples) · 2 pointer-only (licence)
-
Tracking Progress Towards Sustainable Development Goal 6 Using Satellite Imagery 28 Nov 2024 · 0 repositories · arXiv:2411.19093
-
Residual Attention Single-Head Vision Transformer Network for Rolling Bearing Fault Diagnosis in Noisy Environments 27 Nov 2024 · 0 repositories · arXiv:2412.00085
-
MWFormer: Multi-Weather Image Restoration Using Degradation-Aware Transformers 26 Nov 2024 · 1 repository · arXiv:2411.17226
-
SatVision-TOA: A Geospatial Foundation Model for Coarse-Resolution All-Sky Remote Sensing Imagery 26 Nov 2024 · 1 repository · arXiv:2411.17000
-
SCASeg: Strip Cross-Attention for Efficient Semantic Segmentation 26 Nov 2024 · 0 repositories · arXiv:2411.17061
-
CMAViT: Integrating Climate, Managment, and Remote Sensing Data for Crop Yield Estimation with Multimodel Vision Transformers 25 Nov 2024 · 0 repositories · arXiv:2411.16989
-
Factorized Visual Tokenization and Generation 25 Nov 2024 · 0 repositories · arXiv:2411.16681
-
Interpreting Object-level Foundation Models via Visual Precision Search 25 Nov 2024 · 2 repositories · arXiv:2411.16198
-
Soft-TransFormers for Continual Learning 25 Nov 2024 · 1 repository · arXiv:2411.16073
-
UltraSam: A Foundation Model for Ultrasound using Large Open-Access Segmentation Datasets 25 Nov 2024 · 1 repository · arXiv:2411.16222
-
VICON: Vision In-Context Operator Networks for Multi-Physics Fluid Dynamics Prediction 25 Nov 2024 · 1 repository · arXiv:2411.16063
-
Beyond adaptive gradient: Fast-Controlled Minibatch Algorithm for large-scale optimization 24 Nov 2024 · 1 repository · arXiv:2411.15795
-
Evaluating Vision Transformer Models for Visual Quality Control in Industrial Manufacturing 22 Nov 2024 · 1 repository · arXiv:2411.14953
-
Resolution-Agnostic Transformer-based Climate Downscaling 22 Nov 2024 · 0 repositories · arXiv:2411.14774
-
When Spatial meets Temporal in Action Recognition 22 Nov 2024 · 0 repositories · arXiv:2411.15284
-
DINO-X: A Unified Vision Model for Open-World Object Detection and Understanding 21 Nov 2024 · 1 repository · arXiv:2411.14347
-
Quantum Attention for Vision Transformers in High Energy Physics 20 Nov 2024 · 0 repositories · arXiv:2411.13520
-
Faster Multi-GPU Training with PPLL: A Pipeline Parallelism Framework Leveraging Local Learning 19 Nov 2024 · 0 repositories · arXiv:2411.12780
-
Residual Vision Transformer (ResViT) Based Self-Supervised Learning Model for Brain Tumor Classification 19 Nov 2024 · 0 repositories · arXiv:2411.12874
-
DeforHMR: Vision Transformer with Deformable Cross-Attention for 3D Human Mesh Recovery 18 Nov 2024 · 0 repositories · arXiv:2411.11214
-
FCC: Fully Connected Correlation for Few-Shot Segmentation 18 Nov 2024 · 0 repositories · arXiv:2411.11917
-
In-Situ Melt Pool Characterization via Thermal Imaging for Defect Detection in Directed Energy Deposition Using Vision Transformers 18 Nov 2024 · 0 repositories · arXiv:2411.12028
-
MpoxVLM: A Vision-Language Model for Diagnosing Skin Lesions from Mpox Virus Infection 16 Nov 2024 · 1 repository · arXiv:2411.10888
-
A Multi-Scale Spatial-Temporal Network for Wireless Video Transmission 15 Nov 2024 · 0 repositories · arXiv:2411.09936
-
Building 6G Radio Foundation Models with Transformer Architectures 15 Nov 2024 · 0 repositories · arXiv:2411.09996
-
CorrCLIP: Reconstructing Correlations in CLIP with Off-the-Shelf Foundation Models for Open-Vocabulary Semantic Segmentation 15 Nov 2024 · 1 repository · arXiv:2411.10086Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 9 harvested samples) · 9 pointer-only (licence)