Methods › Computer Vision › Vision Transformers › Vision Transformer › Papers, page 5
Vision Transformer
Papers archive 2025-07-28
archive papers tagged: 2,144 · with a code link: 1,051 · where Syntology ran a sample: 328 (286 with a run with no instrument failure, 42 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (328 of 2,144 tagged: 286 with a run with no instrument failure, 42 where every run was a failure of Syntology's instrument)
Page 5 of 22: papers 401 to 500 of 2,144, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Evidential Federated Learning for Skin Lesion Image Classification 15 Nov 2024 · 0 repositories · arXiv:2411.10071
-
Assessing the Performance of the DINOv2 Self-supervised Learning Vision Transformer Model for the Segmentation of the Left Atrium from MRI Images 14 Nov 2024 · 0 repositories · arXiv:2411.09598
-
Harnessing Vision Foundation Models for High-Performance, Training-Free Open Vocabulary Segmentation 14 Nov 2024 · 1 repository · arXiv:2411.09219
-
Learning Parameter Sharing with Tensor Decompositions and Sparsity 14 Nov 2024 · 1 repository · arXiv:2411.09816
-
Partial Multi-View Clustering via Meta-Learning and Contrastive Feature Alignment 14 Nov 2024 · 0 repositories · arXiv:2411.09758
-
SAG-ViT: A Scale-Aware, High-Fidelity Patching Approach with Graph Attention for Vision Transformers 14 Nov 2024 · 1 repository · arXiv:2411.09420
-
AD-DINO: Attention-Dynamic DINO for Distance-Aware Embodied Reference Understanding 13 Nov 2024 · 0 repositories · arXiv:2411.08451
-
DINO-LG: A Task-Specific DINO Model for Coronary Calcium Scoring 12 Nov 2024 · 0 repositories · arXiv:2411.07976
-
ScaleKD: Strong Vision Transformers Could Be Excellent Teachers 11 Nov 2024 · 1 repository · arXiv:2411.06786
-
Track Any Peppers: Weakly Supervised Sweet Pepper Tracking Using VLMs 11 Nov 2024 · 0 repositories · arXiv:2411.06702
-
Few-shot Semantic Learning for Robust Multi-Biome 3D Semantic Mapping in Off-Road Environments 10 Nov 2024 · 0 repositories · arXiv:2411.06632
-
Community Research Earth Digital Intelligence Twin (CREDIT) 9 Nov 2024 · 2 repositories · arXiv:2411.07814Syntology official (archive's flag): 22 ran · 22 ran (of which 0 constructed an object rather than computing a result; 21 with no instrument failure: 1 honoured, 2 violated, 18 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 23 harvested samples) · 2 pointer-only (licence)
-
Pattern Integration and Enhancement Vision Transformer for Self-Supervised Learning in Remote Sensing 9 Nov 2024 · 0 repositories · arXiv:2411.06091
-
ViTOC: Vision Transformer and Object-aware Captioner 9 Nov 2024 · 0 repositories · arXiv:2411.07265
-
Cascaded Dual Vision Transformer for Accurate Facial Landmark Detection 8 Nov 2024 · 1 repository · arXiv:2411.07167
-
Classification of Adventitious Sounds Combining Cochleogram and Vision Transformers 8 Nov 2024 · 0 repositories · arXiv:2411.05955
-
Emotional Images: Assessing Emotions in Images and Potential Biases in Generative Models 8 Nov 2024 · 0 repositories · arXiv:2411.05985
-
GCI-ViTAL: Gradual Confidence Improvement with Vision Transformers for Active Learning on Label Noise 8 Nov 2024 · 0 repositories · arXiv:2411.05939
-
Image inpainting enhancement by replacing the original mask with a self-attended region from the input image 8 Nov 2024 · 0 repositories · arXiv:2411.05705
-
Online-LoRA: Task-free Online Continual Learning via Low Rank Adaptation 8 Nov 2024 · 1 repository · arXiv:2411.05663
-
Tell What You Hear From What You See -- Video to Audio Generation Through Text 8 Nov 2024 · 1 repository · arXiv:2411.05679Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
ViT Enhanced Privacy-Preserving Secure Medical Data Sharing and Classification 8 Nov 2024 · 0 repositories · arXiv:2411.05901
-
DINO-WM: World Models on Pre-trained Visual Features enable Zero-shot Planning 7 Nov 2024 · 0 repositories · arXiv:2411.04983
-
Prion-ViT: Prions-Inspired Vision Transformers for Temperature prediction with Specklegrams 6 Nov 2024 · 0 repositories · arXiv:2411.05836
-
Reducing catastrophic forgetting of incremental learning in the absence of rehearsal memory with task-specific token 6 Nov 2024 · 0 repositories · arXiv:2411.05846
-
LASER: Attention with Exponential Transformation 5 Nov 2024 · 0 repositories · arXiv:2411.03493
-
Optimizing Multi-Scale Representations to Detect Effect Heterogeneity Using Earth Observation and Computer Vision: Applications to Two Anti-Poverty RCTs 4 Nov 2024 · 0 repositories · arXiv:2411.02134
-
V-CAS: A Realtime Vehicle Anti Collision System Using Vision Transformer on Multi-Camera Streams 4 Nov 2024 · 0 repositories · arXiv:2411.01963
-
Aerial Flood Scene Classification Using Fine-Tuned Attention-based Architecture for Flood-Prone Countries in South Asia 31 Oct 2024 · 0 repositories · arXiv:2411.00169
-
Enhancing Brain Tumor Classification Using TrAdaBoost and Multi-Classifier Deep Learning Approaches 31 Oct 2024 · 0 repositories · arXiv:2411.00875
-
JEMA: A Joint Embedding Framework for Scalable Co-Learning with Multimodal Alignment 31 Oct 2024 · 0 repositories · arXiv:2410.23988
-
ViT-LCA: A Neuromorphic Approach for Vision Transformers 31 Oct 2024 · 0 repositories · arXiv:2411.00140
-
Emergence of Human-Like Attention in Self-Supervised Vision Transformers: an eye-tracking study 30 Oct 2024 · 1 repository · arXiv:2410.22768
-
Epipolar-Free 3D Gaussian Splatting for Generalizable Novel View Synthesis 30 Oct 2024 · 0 repositories · arXiv:2410.22817
-
NMformer: A Transformer for Noisy Modulation Classification in Wireless Communication 30 Oct 2024 · 1 repository · arXiv:2411.02428
-
S3PT: Scene Semantics and Structure Guided Clustering to Boost Self-Supervised Pre-Training for Autonomous Driving 30 Oct 2024 · 0 repositories · arXiv:2410.23085
-
DINeuro: Distilling Knowledge from 2D Natural Images via Deformable Tubular Transferring Strategy for 3D Neuron Reconstruction 29 Oct 2024 · 0 repositories · arXiv:2410.22078
-
Multi-step feature fusion for natural disaster damage assessment on satellite images 29 Oct 2024 · 1 repository · arXiv:2410.21901
-
Spatio-temporal Transformers for Action Unit Classification with Event Cameras 29 Oct 2024 · 0 repositories · arXiv:2410.21958
-
Explainability in AI Based Applications: A Framework for Comparing Different Techniques 28 Oct 2024 · 0 repositories · arXiv:2410.20873
-
Interpretable Image Classification with Adaptive Prototype-based Vision Transformers 28 Oct 2024 · 1 repository · arXiv:2410.20722Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 7 harvested samples)
-
Multi-modal AI for comprehensive breast cancer prognostication 28 Oct 2024 · 0 repositories · arXiv:2410.21256
-
Accelerating Augmentation Invariance Pretraining 27 Oct 2024 · 0 repositories · arXiv:2410.22364
-
PViT: Prior-augmented Vision Transformer for Out-of-distribution Detection 27 Oct 2024 · 1 repository · arXiv:2410.20631
-
Enhancing Lie Detection Accuracy: A Comparative Study of Classic ML, CNN, and GCN Models using Audio-Visual Features 26 Oct 2024 · 0 repositories · arXiv:2411.08885
-
Generative Adversarial Patches for Physical Attacks on Cross-Modal Pedestrian Re-Identification 26 Oct 2024 · 0 repositories · arXiv:2410.20097
-
Transforming Precision: A Comparative Analysis of Vision Transformers, CNNs, and Traditional ML for Knee Osteoarthritis Severity Diagnosis 26 Oct 2024 · 0 repositories · arXiv:2410.20062
-
A Multimodal Approach For Endoscopic VCE Image Classification Using BiomedCLIP-PubMedBERT 25 Oct 2024 · 1 repository · arXiv:2410.19944
-
Frozen-DETR: Enhancing DETR with Image Understanding from Frozen Foundation Models 25 Oct 2024 · 0 repositories · arXiv:2410.19635
-
Multi-Class Abnormality Classification Task in Video Capsule Endoscopy 25 Oct 2024 · 1 repository · arXiv:2410.19973
-
FedBaF: Federated Learning Aggregation Biased by a Foundation Model 24 Oct 2024 · 0 repositories · arXiv:2410.18352
-
PESFormer: Boosting Macro- and Micro-expression Spotting with Direct Timestamp Encoding 24 Oct 2024 · 0 repositories · arXiv:2410.18695
-
Multi-scale feature reconstruction network for industrial anomaly detection 23 Oct 2024 · 1 repository
-
DI-MaskDINO: A Joint Object Detection and Instance Segmentation Model 22 Oct 2024 · 1 repository · arXiv:2410.16707
-
Domain-Adaptive Pre-training of Self-Supervised Foundation Models for Medical Image Classification in Gastrointestinal Endoscopy 21 Oct 2024 · 1 repository · arXiv:2410.21302
-
Generalizing Motion Planners with Mixture of Experts for Autonomous Driving 21 Oct 2024 · 1 repository · arXiv:2410.15774Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
ViMoE: An Empirical Study of Designing Vision Mixture-of-Experts 21 Oct 2024 · 0 repositories · arXiv:2410.15732
-
EViT-Unet: U-Net Like Efficient Vision Transformer for Medical Image Segmentation on Mobile and Edge Devices 19 Oct 2024 · 1 repository · arXiv:2410.15036
-
Visual Navigation of Digital Libraries: Retrieval and Classification of Images in the National Library of Norway's Digitised Book Collection 19 Oct 2024 · 1 repository · arXiv:2410.14969
-
LUDVIG: Learning-free Uplifting of 2D Visual features to Gaussian Splatting scenes 18 Oct 2024 · 0 repositories · arXiv:2410.14462
-
Co-Segmentation without any Pixel-level Supervision with Application to Large-Scale Sketch Classification 17 Oct 2024 · 0 repositories · arXiv:2410.13582
-
On Partial Prototype Collapse in the DINO Family of Self-Supervised Methods 17 Oct 2024 · 0 repositories · arXiv:2410.14060
-
Training Compute-Optimal Vision Transformers for Brain Encoding 17 Oct 2024 · 0 repositories · arXiv:2410.19810
-
Efficient Partitioning Vision Transformer on Edge Devices for Distributed Inference 15 Oct 2024 · 0 repositories · arXiv:2410.11650
-
Pixology: Probing the Linguistic and Visual Capabilities of Pixel-based Language Models 15 Oct 2024 · 1 repository · arXiv:2410.12011Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 8 harvested samples) · 8 pointer-only (licence)
-
Unveiling the Mystery of Visual Attributes of Concrete and Abstract Concepts: Variability, Nearest Neighbors, and Challenging Categories 15 Oct 2024 · 1 repository · arXiv:2410.11657
-
Visual Fixation-Based Retinal Prosthetic Simulation 15 Oct 2024 · 0 repositories · arXiv:2410.11688
-
big.LITTLE Vision Transformer for Efficient Visual Recognition 14 Oct 2024 · 0 repositories · arXiv:2410.10267
-
Performance Evaluation of Deep Learning and Transformer Models Using Multimodal Data for Breast Cancer Classification 14 Oct 2024 · 0 repositories · arXiv:2410.10146
-
Data Adaptive Few-shot Multi Label Segmentation with Foundation Model 13 Oct 2024 · 0 repositories · arXiv:2410.09759
-
Token Pruning using a Lightweight Background Aware Vision Transformer 12 Oct 2024 · 0 repositories · arXiv:2410.09324
-
DeBiFormer: Vision Transformer with Deformable Agent Bi-level Routing Attention 11 Oct 2024 · 1 repository · arXiv:2410.08582
-
ViT3D Alignment of LLaMA3: 3D Medical Image Report Generation 11 Oct 2024 · 0 repositories · arXiv:2410.08588
-
IceDiff: High Resolution and High-Quality Sea Ice Forecasting with Generative Diffusion Prior 10 Oct 2024 · 0 repositories · arXiv:2410.09111
-
SPA: 3D Spatial-Awareness Enables Effective Embodied Representation 10 Oct 2024 · 1 repository · arXiv:2410.08208Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 6 harvested samples)
-
Bridge the Points: Graph-based Few-shot Segment Anything Semantically 9 Oct 2024 · 1 repository · arXiv:2410.06964Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 4 harvested samples)
-
Pair-VPR: Place-Aware Pre-training and Contrastive Pair Classification for Visual Place Recognition with Vision Transformers 9 Oct 2024 · 1 repository · arXiv:2410.06614
-
IncSAR: A Dual Fusion Incremental Learning Framework for SAR Target Recognition 8 Oct 2024 · 1 repository · arXiv:2410.05820
-
Tackling the Abstraction and Reasoning Corpus with Vision Transformers: the Importance of 2D Representation, Positions, and Objects 8 Oct 2024 · 1 repository · arXiv:2410.06405Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples)
-
Vision Transformer based Random Walk for Group Re-Identification 8 Oct 2024 · 0 repositories · arXiv:2410.05808
-
Improving Image Clustering with Artifacts Attenuation via Inference-Time Attention Engineering 7 Oct 2024 · 0 repositories · arXiv:2410.04801
-
Low-Rank Continual Pyramid Vision Transformer: Incrementally Segment Whole-Body Organs in CT with Light-Weighted Adaptation 7 Oct 2024 · 0 repositories · arXiv:2410.04689
-
Optimizing Medical Image Segmentation with Advanced Decoder Design 5 Oct 2024 · 1 repository · arXiv:2410.04128
-
Self-Supervised Anomaly Detection in the Wild: Favor Joint Embeddings Methods 5 Oct 2024 · 0 repositories · arXiv:2410.04289
-
An X-Ray Is Worth 15 Features: Sparse Autoencoders for Interpretable Radiology Report Generation 4 Oct 2024 · 0 repositories · arXiv:2410.03334
-
HiFiSeg: High-Frequency Information Enhanced Polyp Segmentation with Global-Local Vision Transformer 3 Oct 2024 · 0 repositories · arXiv:2410.02528
-
A versatile machine learning workflow for high-throughput analysis of supported metal catalyst particles 2 Oct 2024 · 1 repository · arXiv:2410.01213
-
Depth Pro: Sharp Monocular Metric Depth in Less Than a Second 2 Oct 2024 · 1 repository · arXiv:2410.02073Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
Deep Multimodal Fusion for Semantic Segmentation of Remote Sensing Earth Observation Data 1 Oct 2024 · 0 repositories · arXiv:2410.00469
-
Towards Open-Vocabulary Semantic Segmentation Without Semantic Labels 30 Sep 2024 · 0 repositories · arXiv:2409.19846
-
Discerning the Chaos: Detecting Adversarial Perturbations while Disentangling Intentional from Unintentional Noises 29 Sep 2024 · 0 repositories · arXiv:2409.19619
-
Unveil Benign Overfitting for Transformer in Vision: Training Dynamics, Convergence, and Generalization 28 Sep 2024 · 0 repositories · arXiv:2409.19345
-
How Effective is Pre-training of Large Masked Autoencoders for Downstream Earth Observation Tasks? 27 Sep 2024 · 0 repositories · arXiv:2409.18536
-
Improving Visual Object Tracking through Visual Prompting 27 Sep 2024 · 1 repository · arXiv:2409.18901
-
Developing a Dual-Stage Vision Transformer Model for Lung Disease Classification 26 Sep 2024 · 0 repositories · arXiv:2409.18257
-
Ophthalmic Biomarker Detection with Parallel Prediction of Transformer and Convolutional Architecture 26 Sep 2024 · 0 repositories · arXiv:2409.17788
-
Self-supervised Pretraining for Cardiovascular Magnetic Resonance Cine Segmentation 26 Sep 2024 · 1 repository · arXiv:2409.18100
-
Block Expanded DINORET: Adapting Natural Domain Foundation Models for Retinal Imaging Without Catastrophic Forgetting 25 Sep 2024 · 0 repositories · arXiv:2409.17332
-
HVT: A Comprehensive Vision Framework for Learning in Non-Euclidean Space 25 Sep 2024 · 1 repository · arXiv:2409.16897
-
Clinical-grade Multi-Organ Pathology Report Generation for Multi-scale Whole Slide Images via a Semantically Guided Medical Text Foundation Model 23 Sep 2024 · 1 repository · arXiv:2409.15574