Methods › Computer Vision › Vision Transformers › Vision Transformer › Papers, page 3
Vision Transformer
Papers archive 2025-07-28
archive papers tagged: 2,144 · with a code link: 1,051 · where Syntology ran a sample: 328 (286 with a run with no instrument failure, 42 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (328 of 2,144 tagged: 286 with a run with no instrument failure, 42 where every run was a failure of Syntology's instrument)
Page 3 of 22: papers 201 to 300 of 2,144, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Developing a PET/CT Foundation Model for Cross-Modal Anatomical and Functional Imaging 4 Mar 2025 · 0 repositories · arXiv:2503.02824
-
TeTRA-VPR: A Ternary Transformer Approach for Compact Visual Place Recognition 4 Mar 2025 · 0 repositories · arXiv:2503.02511
-
Label Ranker: Self-Aware Preference for Classification Label Position in Visual Masked Self-Supervised Pre-Trained Model 3 Mar 2025 · 1 repository
-
MI-DETR: An Object Detection Model with Multi-time Inquiries Mechanism 3 Mar 2025 · 1 repository · arXiv:2503.01463
-
MRI super-resolution reconstruction using efficient diffusion probabilistic model with residual shifting 3 Mar 2025 · 1 repository · arXiv:2503.01576
-
ViKANformer: Embedding Kolmogorov Arnold Networks in Vision Transformers for Pattern-Based Learning 3 Mar 2025 · 0 repositories · arXiv:2503.01124
-
An Integrated Deep Learning Framework Leveraging NASNet and Vision Transformer with MixProcessing for Accurate and Precise Diagnosis of Lung Diseases 27 Feb 2025 · 0 repositories · arXiv:2502.20570
-
Do computer vision foundation models learn the low-level characteristics of the human visual system? 27 Feb 2025 · 0 repositories · arXiv:2502.20256
-
Regional climate projections using a deep-learning-based model-ranking and downscaling framework: Application to European climate zones 27 Feb 2025 · 0 repositories · arXiv:2502.20132
-
Revisit the Stability of Vanilla Federated Learning Under Diverse Conditions 27 Feb 2025 · 0 repositories · arXiv:2502.19849
-
WalnutData: A UAV Remote Sensing Dataset of Green Walnuts and Model Evaluation 27 Feb 2025 · 1 repository · arXiv:2502.20092
-
Brain-inspired analogical mixture prototypes for few-shot class-incremental learning 26 Feb 2025 · 0 repositories · arXiv:2502.18923
-
Examining the Threat Landscape: Foundation Models and Model Stealing 25 Feb 2025 · 0 repositories · arXiv:2502.18077
-
CalibRefine: Deep Learning-Based Online Automatic Targetless LiDAR-Camera Calibration with Iterative and Attention-Driven Post-Refinement 24 Feb 2025 · 1 repository · arXiv:2502.17648
-
ENACT-Heart -- ENsemble-based Assessment Using CNN and Transformer on Heart Sounds 24 Feb 2025 · 0 repositories · arXiv:2502.16914
-
Enhancing Image Matting in Real-World Scenes with Mask-Guided Iterative Refinement 24 Feb 2025 · 0 repositories · arXiv:2502.17093
-
MaxGlaViT: A novel lightweight vision transformer-based approach for early diagnosis of glaucoma stages from fundus images 24 Feb 2025 · 1 repository · arXiv:2502.17154
-
Unraveling the geometry of visual relational reasoning 24 Feb 2025 · 1 repository · arXiv:2502.17382
-
VPNeXt -- Rethinking Dense Decoding for Plain Vision Transformer 23 Feb 2025 · 0 repositories · arXiv:2502.16654
-
Vision Transformer Accelerator ASIC for Real-Time, Low-Power Sleep Staging 22 Feb 2025 · 0 repositories · arXiv:2502.16334
-
Mantis: Lightweight Calibrated Foundation Model for User-Friendly Time Series Classification 21 Feb 2025 · 1 repository · arXiv:2502.15637Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples)
-
Medical Image Classification with KAN-Integrated Transformers and Dilated Neighborhood Attention 19 Feb 2025 · 1 repository · arXiv:2502.13693
-
Qwen2.5-VL Technical Report 19 Feb 2025 · 4 repositories · arXiv:2502.13923Syntology 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples)
-
Myna: Masking-Based Contrastive Learning of Musical Representations 18 Feb 2025 · 1 repository · arXiv:2502.12511
-
Towards an automated workflow in materials science for combining multi-modal simulative and experimental information using data mining and large language models 18 Feb 2025 · 0 repositories · arXiv:2502.14904
-
OCT Data is All You Need: How Vision Transformers with and without Pre-training Benefit Imaging 17 Feb 2025 · 0 repositories · arXiv:2502.12379
-
A recurrent vision transformer shows signatures of primate visual attention 16 Feb 2025 · 0 repositories · arXiv:2502.10955
-
Automatic Quality Assessment of First Trimester Crown-Rump-Length Ultrasound Images 15 Feb 2025 · 0 repositories · arXiv:2502.10908
-
CLoCKDistill: Consistent Location-and-Context-aware Knowledge Distillation for DETRs 15 Feb 2025 · 0 repositories · arXiv:2502.10683
-
A synergistic CNN-transformer network with pooling attention fusion for hyperspectral image classification 14 Feb 2025 · 1 repository
-
Compress image to patches for Vision Transformer 14 Feb 2025 · 1 repository · arXiv:2502.10120
-
Janus: Collaborative Vision Transformer Under Dynamic Network Environment 14 Feb 2025 · 0 repositories · arXiv:2502.10047
-
QMaxViT-Unet+: A Query-Based MaxViT-Unet with Edge Enhancement for Scribble-Supervised Segmentation of Medical Images 14 Feb 2025 · 1 repository · arXiv:2502.10294
-
Simplifying DINO via Coding Rate Regularization 14 Feb 2025 · 0 repositories · arXiv:2502.10385
-
Hierarchical Vision Transformer with Prototypes for Interpretable Medical Image Classification 13 Feb 2025 · 0 repositories · arXiv:2502.08997
-
Residual Transformer Fusion Network for Salt and Pepper Image Denoising 13 Feb 2025 · 0 repositories · arXiv:2502.09000
-
Hi-End-MAE: Hierarchical encoder-driven masked autoencoders are stronger vision learners for medical image segmentation 12 Feb 2025 · 1 repository · arXiv:2502.08347
-
5D Neural Surrogates for Nonlinear Gyrokinetic Simulations of Plasma Turbulence 11 Feb 2025 · 0 repositories · arXiv:2502.07469
-
Dataset Ownership Verification in Contrastive Pre-trained Models 11 Feb 2025 · 1 repository · arXiv:2502.07276
-
Fast-COS: A Fast One-Stage Object Detector Based on Reparameterized Attention Vision Transformer for Autonomous Driving 11 Feb 2025 · 0 repositories · arXiv:2502.07417
-
Fully Exploiting Vision Foundation Model's Profound Prior Knowledge for Generalizable RGB-Depth Driving Scene Parsing 10 Feb 2025 · 0 repositories · arXiv:2502.06219
-
Multimodal Task Representation Memory Bank vs. Catastrophic Forgetting in Anomaly Detection 10 Feb 2025 · 0 repositories · arXiv:2502.06194
-
Unconstrained Body Recognition at Altitude and Range: Comparing Four Approaches 10 Feb 2025 · 0 repositories · arXiv:2502.07130
-
ViSIR: Vision Transformer Single Image Reconstruction Method for Earth System Models 10 Feb 2025 · 0 repositories · arXiv:2502.06741
-
Topological derivative approach for deep neural network architecture adaptation 8 Feb 2025 · 0 repositories · arXiv:2502.06885
-
MedMimic: Physician-Inspired Multimodal Fusion for Early Diagnosis of Fever of Unknown Origin 7 Feb 2025 · 0 repositories · arXiv:2502.04794
-
SelaFD:Seamless Adaptation of Vision Transformer Fine-tuning for Radar-based Human Activity 7 Feb 2025 · 1 repository · arXiv:2502.04740
-
A Self-supervised Multimodal Deep Learning Approach to Differentiate Post-radiotherapy Progression from Pseudoprogression in Glioblastoma 6 Feb 2025 · 0 repositories · arXiv:2502.03999
-
Vision-Integrated LLMs for Autonomous Driving Assistance : Human Performance Comparison and Trust Evaluation 6 Feb 2025 · 0 repositories · arXiv:2502.06843
-
Label Anything: An Interpretable, High-Fidelity and Prompt-Free Annotator 5 Feb 2025 · 0 repositories · arXiv:2502.02972
-
Maximizing the Position Embedding for Vision Transformers with Global Average Pooling 5 Feb 2025 · 0 repositories · arXiv:2502.02919
-
Optimizing Robustness and Accuracy in Mixture of Experts: A Dual-Model Approach 5 Feb 2025 · 0 repositories · arXiv:2502.06832
-
ZISVFM: Zero-Shot Object Instance Segmentation in Indoor Robotic Environments with Vision Foundation Models 5 Feb 2025 · 0 repositories · arXiv:2502.03266
-
Memory Efficient Transformer Adapter for Dense Predictions 4 Feb 2025 · 0 repositories · arXiv:2502.01962
-
Mind the Gap: Evaluating Patch Embeddings from General-Purpose and Histopathology Foundation Models for Cell Segmentation and Classification 4 Feb 2025 · 1 repository · arXiv:2502.02471
-
The Skin Game: Revolutionizing Standards for AI Dermatology Model Comparison 4 Feb 2025 · 1 repository · arXiv:2502.02500
-
UniGaze: Towards Universal Gaze Estimation via Large-scale Pre-Training 4 Feb 2025 · 0 repositories · arXiv:2502.02307
-
A framework for river connectivity classification using temporal image processing and attention based neural networks 1 Feb 2025 · 0 repositories · arXiv:2502.00474
-
A Study on the Performance of U-Net Modifications in Retroperitoneal Tumor Segmentation 1 Feb 2025 · 1 repository · arXiv:2502.00314
-
Contrastive Forward-Forward: A Training Algorithm of Vision Transformer 1 Feb 2025 · 0 repositories · arXiv:2502.00571
-
CerraData-4MM: A multimodal benchmark dataset on Cerrado for land use and land cover classification 31 Jan 2025 · 1 repository · arXiv:2502.00083
-
From Semantic Segmentation of Natural Images to Medical Image Segmentation Using ViT-Based Architectures 31 Jan 2025 · 0 repositories
-
PixelWorld: Towards Perceiving Everything as Pixels 31 Jan 2025 · 0 repositories · arXiv:2501.19339
-
Arbitrary Data as Images: Fusion of Patient Data Across Modalities and Irregular Intervals with Vision Transformers 30 Jan 2025 · 0 repositories · arXiv:2501.18237
-
Self-Supervised Frameworks for Speaker Verification via Bootstrapped Positive Sampling 29 Jan 2025 · 1 repository · arXiv:2501.17772
-
TransRAD: Retentive Vision Transformer for Enhanced Radar Object Detection 29 Jan 2025 · 1 repository · arXiv:2501.17977
-
Watch Your STEPP: Semantic Traversability Estimation using Pose Projected Features 29 Jan 2025 · 0 repositories · arXiv:2501.17594
-
An Attention-Locating Algorithm for Eliminating Background Effects in Fine-grained Visual Classification 28 Jan 2025 · 1 repository
-
ViT-2SPN: Vision Transformer-based Dual-Stream Self-Supervised Pretraining Networks for Retinal OCT Classification 28 Jan 2025 · 1 repository · arXiv:2501.17260
-
Cross-Domain Semantic Segmentation with Large Language Model-Assisted Descriptor Generation 27 Jan 2025 · 0 repositories · arXiv:2501.16467
-
Leveraging Video Vision Transformer for Alzheimer's Disease Diagnosis from 3D Brain MRI 27 Jan 2025 · 0 repositories · arXiv:2501.15733
-
PDC-ViT : Source Camera Identification using Pixel Difference Convolution and Vision Transformer 27 Jan 2025 · 0 repositories · arXiv:2501.16227
-
AI-Driven Secure Data Sharing: A Trustworthy and Privacy-Preserving Approach 26 Jan 2025 · 0 repositories · arXiv:2501.15363
-
Self-supervised Benchmark Lottery on ImageNet: Do Marginal Improvements Translate to Improvements on Similar Datasets? 26 Jan 2025 · 0 repositories · arXiv:2501.15431
-
Automatic detection and prediction of nAMD activity change in retinal OCT using Siamese networks and Wasserstein Distance for ordinality 24 Jan 2025 · 1 repository · arXiv:2501.14323
-
Surface Vision Mamba: Leveraging Bidirectional State Space Model for Efficient Spherical Manifold Representation 24 Jan 2025 · 0 repositories · arXiv:2501.14679
-
Ensuring Medical AI Safety: Explainable AI-Driven Detection and Mitigation of Spurious Model Behavior and Associated Data 23 Jan 2025 · 1 repository · arXiv:2501.13818
-
Multimodal AI on Wound Images and Clinical Notes for Home Patient Referral 22 Jan 2025 · 0 repositories · arXiv:2501.13247
-
Unified CNNs and transformers underlying learning mechanism reveals multi-head attention modus vivendi 22 Jan 2025 · 0 repositories · arXiv:2501.12900
-
Efficient Lung Ultrasound Severity Scoring Using Dedicated Feature Extractor 21 Jan 2025 · 1 repository · arXiv:2501.12524
-
Vision-Language Models for Automated Chest X-ray Interpretation: Leveraging ViT and GPT-2 21 Jan 2025 · 0 repositories · arXiv:2501.12356
-
Generative AI-enabled Blockage Prediction for Robust Dual-Band mmWave Communication 20 Jan 2025 · 0 repositories · arXiv:2501.11763
-
Efficient Auto-Labeling of Large-Scale Poultry Datasets (ALPD) Using Semi-Supervised Models, Active Learning, and Prompt-then-Detect Approach 18 Jan 2025 · 0 repositories · arXiv:2501.10809
-
Enhancing the Reliability in Machine Learning for Gravitational Wave Parameter Estimation with Attention-Based Models 17 Jan 2025 · 0 repositories · arXiv:2501.10486
-
FiLo++: Zero-/Few-Shot Anomaly Detection by Fused Fine-Grained Descriptions and Deformable Localization 17 Jan 2025 · 1 repository · arXiv:2501.10067
-
Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation 16 Jan 2025 · 1 repository · arXiv:2501.09688Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 1 honoured, 0 violated, 7 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 11 harvested samples) · 3 pointer-only (licence)
-
Generalized Single-Image-Based Morphing Attack Detection Using Deep Representations from Vision Transformer 16 Jan 2025 · 0 repositories · arXiv:2501.09817
-
Learnings from Scaling Visual Tokenizers for Reconstruction and Generation 16 Jan 2025 · 0 repositories · arXiv:2501.09755
-
Prompt-CAM: A Simpler Interpretable Transformer for Fine-Grained Analysis 16 Jan 2025 · 1 repository · arXiv:2501.09333Syntology 1 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified; the one sample that ran constructed an object rather than computing a result (of 1 harvested sample)
-
Deep Self-Supervised Disturbance Mapping with the OPERA Sentinel-1 Radiometric Terrain Corrected SAR Backscatter Product 15 Jan 2025 · 1 repository · arXiv:2501.09129
-
MIAFEx: An Attention-based Feature Extraction Method for Medical Image Classification 15 Jan 2025 · 0 repositories · arXiv:2501.08562
-
SuperSAM: Crafting a SAM Supernetwork via Structured Pruning and Unstructured Parameter Prioritization 15 Jan 2025 · 1 repository · arXiv:2501.08504
-
SwinTExCo: Exemplar-based video colorization using Swin Transformer 15 Jan 2025 · 1 repository
-
Efficient Deep Learning-based Forward Solvers for Brain Tumor Growth Models 14 Jan 2025 · 1 repository · arXiv:2501.08226
-
SST-EM: Advanced Metrics for Evaluating Semantic, Spatial and Temporal Aspects in Video Editing 13 Jan 2025 · 1 repository · arXiv:2501.07554
-
Transforming Vision Transformer: Towards Efficient Multi-Task Asynchronous Learning 12 Jan 2025 · 1 repository · arXiv:2501.06884Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 4 unverified (of 7 harvested samples) · 7 pointer-only (licence)
-
FocusDD: Real-World Scene Infusion for Robust Dataset Distillation 11 Jan 2025 · 0 repositories · arXiv:2501.06405
-
A Holistically Point-guided Text Framework for Weakly-Supervised Camouflaged Object Detection 10 Jan 2025 · 0 repositories · arXiv:2501.06038
-
An Attention-Guided Deep Learning Approach for Classifying 39 Skin Lesion Types 10 Jan 2025 · 1 repository · arXiv:2501.05991
-
Merging Feed-Forward Sublayers for Compressed Transformers 10 Jan 2025 · 1 repository · arXiv:2501.06126