Methods › Computer Vision › Vision Transformers › Vision Transformer › Papers, page 6
Vision Transformer
Papers archive 2025-07-28
archive papers tagged: 2,144 · with a code link: 1,051 · where Syntology ran a sample: 328 (286 with a run with no instrument failure, 42 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (328 of 2,144 tagged: 286 with a run with no instrument failure, 42 where every run was a failure of Syntology's instrument)
Page 6 of 22: papers 501 to 600 of 2,144, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Designing Pre-training Datasets from Unlabeled Data for EEG Classification with Transformers 23 Sep 2024 · 0 repositories · arXiv:2410.07190
-
HydroVision: LiDAR-Guided Hydrometric Prediction with Vision Transformers and Hybrid Graph Learning 23 Sep 2024 · 0 repositories · arXiv:2409.15213
-
Patch Ranking: Efficient CLIP by Learning to Rank Local Patches 22 Sep 2024 · 1 repository · arXiv:2409.14607
-
Multiple-Exit Tuning: Towards Inference-Efficient Adaptation for Vision Transformer 21 Sep 2024 · 0 repositories · arXiv:2409.13999
-
Tackling fluffy clouds: field boundaries detection using time series of S2 and/or S1 imagery 20 Sep 2024 · 1 repository · arXiv:2409.13568
-
ViTGuard: Attention-aware Detection against Adversarial Examples for Vision Transformer 20 Sep 2024 · 0 repositories · arXiv:2409.13828
-
Bridging Domain Gap for Flight-Ready Spaceborne Vision 18 Sep 2024 · 0 repositories · arXiv:2409.11661
-
NT-ViT: Neural Transcoding Vision Transformers for EEG-to-fMRI Synthesis 18 Sep 2024 · 0 repositories · arXiv:2409.11836
-
Unsupervised Feature Orthogonalization for Learning Distortion-Invariant Representations 18 Sep 2024 · 1 repository · arXiv:2409.12276
-
Are Deep Learning Models Robust to Partial Object Occlusion in Visual Recognition Tasks? 16 Sep 2024 · 0 repositories · arXiv:2409.10775
-
MetaFormer and CNN Hybrid Model for Polyp Image Segmentation 16 Sep 2024 · 1 repository
-
Personalized Speech Emotion Recognition in Human-Robot Interaction using Vision Transformers 16 Sep 2024 · 0 repositories · arXiv:2409.10687
-
Underwater Image Enhancement via Dehazing and Color Restoration 15 Sep 2024 · 0 repositories · arXiv:2409.09779
-
Investigation of Hierarchical Spectral Vision Transformer Architecture for Classification of Hyperspectral Imagery 14 Sep 2024 · 0 repositories · arXiv:2409.09244
-
SEA-ViT: Sea Surface Currents Forecasting Using Vision Transformer and GRU-Based Spatio-Temporal Covariance Modeling 14 Sep 2024 · 1 repository · arXiv:2409.16313
-
HTR-VT: Handwritten Text Recognition with Vision Transformer 13 Sep 2024 · 2 repositories · arXiv:2409.08573
-
Pathfinder for Low-altitude Aircraft with Binary Neural Network 13 Sep 2024 · 1 repository · arXiv:2409.08824
-
Phikon-v2, A large and public feature extractor for biomarker prediction 13 Sep 2024 · 0 repositories · arXiv:2409.09173
-
AD-Lite Net: A Lightweight and Concatenated CNN Model for Alzheimer's Detection from MRI Images 12 Sep 2024 · 0 repositories · arXiv:2409.08170
-
Intrapartum Ultrasound Image Segmentation of Pubic Symphysis and Fetal Head Using Dual Student-Teacher Framework with CNN-ViT Collaborative Learning 11 Sep 2024 · 1 repository · arXiv:2409.06928
-
Token Turing Machines are Efficient Vision Models 11 Sep 2024 · 1 repository · arXiv:2409.07613Syntology official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 11 with no instrument failure: 1 honoured, 0 violated, 10 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 14 harvested samples) · 1 pointer-only (licence)
-
Connecting Concept Convexity and Human-Machine Alignment in Deep Neural Networks 10 Sep 2024 · 0 repositories · arXiv:2409.06362
-
Static for Dynamic: Towards a Deeper Understanding of Dynamic Facial Expressions Using Static Expression Data 10 Sep 2024 · 1 repository · arXiv:2409.06154
-
Exploring Rich Subjective Quality Information for Image Quality Assessment in the Wild 9 Sep 2024 · 0 repositories · arXiv:2409.05540
-
Sequential Posterior Sampling with Diffusion Models 9 Sep 2024 · 0 repositories · arXiv:2409.05399
-
LMLT: Low-to-high Multi-Level Vision Transformer for Image Super-Resolution 5 Sep 2024 · 1 repository · arXiv:2409.03516
-
Onboard Satellite Image Classification for Earth Observation: A Comparative Study of ViT Models 5 Sep 2024 · 1 repository · arXiv:2409.03901
-
Towards Data-Centric Face Anti-Spoofing: Improving Cross-domain Generalization via Physics-based Data Synthesis 4 Sep 2024 · 0 repositories · arXiv:2409.03501
-
AstroMAE: Redshift Prediction Using a Masked Autoencoder with a Novel Fine-Tuning Architecture 3 Sep 2024 · 0 repositories · arXiv:2409.01825
-
Improving Apple Object Detection with Occlusion-Enhanced Distillation 3 Sep 2024 · 0 repositories · arXiv:2409.01573
-
T1-contrast Enhanced MRI Generation from Multi-parametric MRI for Glioma Patients with Latent Tumor Conditioning 3 Sep 2024 · 0 repositories · arXiv:2409.01622
-
Evidential Transformers for Improved Image Retrieval 2 Sep 2024 · 0 repositories · arXiv:2409.01082
-
MVX-ViT: Multimodal Collaborative Perception for 6G V2X Network Management Decisions Using Vision Transformer. 30 Aug 2024 · 1 repository
-
Evaluating Deep Learning Models for Breast Cancer Classification: A Comparative Study 29 Aug 2024 · 2 repositories · arXiv:2408.16859
-
LLaVA-SG: Leveraging Scene Graphs as Visual Semantic Expression in Vision-Language Models 29 Aug 2024 · 0 repositories · arXiv:2408.16224
-
PartFormer: Awakening Latent Diverse Representation from Vision Transformer for Object Re-Identification 29 Aug 2024 · 0 repositories · arXiv:2408.16684
-
Weakly Supervised Object Detection for Automatic Tooth-marked Tongue Recognition 29 Aug 2024 · 1 repository · arXiv:2408.16451
-
Applying ViT in Generalized Few-shot Semantic Segmentation 27 Aug 2024 · 1 repository · arXiv:2408.14957
-
The Benefits of Balance: From Information Projections to Variance Reduction 27 Aug 2024 · 0 repositories · arXiv:2408.15065
-
3D-RCNet: Learning from Transformer to Build a 3D Relational ConvNet for Hyperspectral Image Classification 25 Aug 2024 · 1 repository · arXiv:2408.13728
-
AlphaViT: A Flexible Game-Playing AI for Multiple Games and Variable Board Sizes 25 Aug 2024 · 1 repository · arXiv:2408.13871
-
LowCLIP: Adapting the CLIP Model Architecture for Low-Resource Languages in Multimodal Image Retrieval Task 25 Aug 2024 · 0 repositories · arXiv:2408.13909
-
Variational Autoencoder for Anomaly Detection: A Comparative Study 24 Aug 2024 · 1 repository · arXiv:2408.13561
-
EAViT: External Attention Vision Transformer for Audio Classification 23 Aug 2024 · 0 repositories · arXiv:2408.13201
-
Image Segmentation in Foundation Model Era: A Survey 23 Aug 2024 · 1 repository · arXiv:2408.12957
-
Enhanced Infield Agriculture with Interpretable Machine Learning Approaches for Crop Classification 22 Aug 2024 · 0 repositories · arXiv:2408.12426
-
Sapiens: Foundation for Human Vision Models 22 Aug 2024 · 2 repositories · arXiv:2408.12569Syntology 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 11 harvested samples)
-
On Learnable Parameters of Optimal and Suboptimal Deep Learning Models 21 Aug 2024 · 0 repositories · arXiv:2408.11720
-
Toward Enhancing Vehicle Color Recognition in Adverse Conditions: A Dataset and Benchmark 21 Aug 2024 · 1 repository · arXiv:2408.11589
-
HiRED: Attention-Guided Token Dropping for Efficient Inference of High-Resolution Vision-Language Models 20 Aug 2024 · 1 repository · arXiv:2408.10945Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 3 harvested samples)
-
MambaEVT: Event Stream based Visual Object Tracking using State Space Model 20 Aug 2024 · 1 repository · arXiv:2408.10487Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 1 honoured, 0 violated, 8 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 10 harvested samples)
-
On the Potential of Open-Vocabulary Models for Object Detection in Unusual Street Scenes 20 Aug 2024 · 0 repositories · arXiv:2408.11221
-
Quantum Inverse Contextual Vision Transformers (Q-ICVT): A New Frontier in 3D Object Detection for AVs 20 Aug 2024 · 1 repository · arXiv:2408.11207
-
Pedestrian Attribute Recognition: A New Benchmark Dataset and A Large Language Model Augmented Framework 19 Aug 2024 · 2 repositories · arXiv:2408.09720
-
SAM-UNet:Enhancing Zero-Shot Segmentation of SAM for Universal Medical Images 19 Aug 2024 · 1 repository · arXiv:2408.09886
-
OU-CoViT: Copula-Enhanced Bi-Channel Multi-Task Vision Transformers with Dual Adaptation for OU-UWF Images 18 Aug 2024 · 0 repositories · arXiv:2408.09395
-
A Novel Approach to Classify Power Quality Signals Using Vision Transformers 16 Aug 2024 · 0 repositories · arXiv:2409.00025
-
Comparative Analysis of Generative Models: Enhancing Image Synthesis with VAEs, GANs, and Stable Diffusion 16 Aug 2024 · 0 repositories · arXiv:2408.08751
-
Research on Personalized Compression Algorithm for Pre-trained Models Based on Homomorphic Entropy Increase 16 Aug 2024 · 0 repositories · arXiv:2408.08684
-
Computer Vision Model Compression Techniques for Embedded Systems: A Survey 15 Aug 2024 · 1 repository · arXiv:2408.08250
-
Distributional Drift Detection in Medical Imaging with Sketching and Fine-Tuned Transformer 15 Aug 2024 · 0 repositories · arXiv:2408.08456
-
Unsupervised Part Discovery via Dual Representation Alignment 15 Aug 2024 · 1 repository · arXiv:2408.08108
-
G²V²former: Graph Guided Video Vision Transformer for Face Anti-Spoofing 14 Aug 2024 · 0 repositories · arXiv:2408.07675
-
Cross-View Geolocalization and Disaster Mapping with Street-View and VHR Satellite Imagery: A Case Study of Hurricane IAN 13 Aug 2024 · 1 repository · arXiv:2408.06761
-
PIR: Photometric Inverse Rendering with Shading Cues Modeling and Surface Reflectance Regularization 13 Aug 2024 · 1 repository · arXiv:2408.06828
-
ReCLIP++: Learn to Rectify the Bias of CLIP for Unsupervised Semantic Segmentation 13 Aug 2024 · 1 repository · arXiv:2408.06747Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
Spectrum Prediction With Deep 3D Pyramid Vision Transformer Learning 13 Aug 2024 · 1 repository · arXiv:2408.06870
-
An analysis of HOI: using a training-free method with multimodal visual foundation models when only the test set is available, without the training set 11 Aug 2024 · 0 repositories · arXiv:2408.05772
-
U-DECN: End-to-End Underwater Object Detection ConvNet with Improved DeNoising Training 11 Aug 2024 · 1 repository · arXiv:2408.05780
-
BeyondCT: A deep learning model for predicting pulmonary function from chest CT scans 10 Aug 2024 · 0 repositories · arXiv:2408.05645
-
PersonViT: Large-scale Self-supervised Vision Transformer for Person Re-Identification 10 Aug 2024 · 1 repository · arXiv:2408.05398
-
BRAT: Bonus oRthogonAl Token for Architecture Agnostic Textual Inversion 8 Aug 2024 · 1 repository · arXiv:2408.04785
-
M2EF-NNs: Multimodal Multi-instance Evidence Fusion Neural Networks for Cancer Survival Prediction 8 Aug 2024 · 0 repositories · arXiv:2408.04170
-
UHNet: An Ultra-Lightweight and High-Speed Edge Detection Network 8 Aug 2024 · 0 repositories · arXiv:2408.04258
-
No-Reference Image Quality Assessment with Global-Local Progressive Integration and Semantic-Aligned Quality Transfer 7 Aug 2024 · 1 repository · arXiv:2408.03885
-
RailTrack-DaViT: A Vision Transformer-Based Approach for Automated Railway Track Defect Detection 7 Aug 2024 · 1 repository
-
Advancing EEG-Based Gaze Prediction Using Depthwise Separable Convolution and Enhanced Pre-Processing 6 Aug 2024 · 1 repository · arXiv:2408.03480
-
AssemAI: Interpretable Image-Based Anomaly Detection for Manufacturing Pipelines 5 Aug 2024 · 1 repository · arXiv:2408.02181
-
LAM3D: Leveraging Attention for Monocular 3D Object Detection 3 Aug 2024 · 0 repositories · arXiv:2408.01739
-
Privacy-Preserving Split Learning with Vision Transformers using Patch-Wise Random and Noisy CutMix 2 Aug 2024 · 0 repositories · arXiv:2408.01040
-
THOR2: Topological Analysis for 3D Shape and Color-Based Human-Inspired Object Recognition in Unseen Environments 2 Aug 2024 · 1 repository · arXiv:2408.01579
-
CC-SAM: SAM with Cross-feature Attention and Context for Ultrasound Image Segmentation 31 Jul 2024 · 0 repositories · arXiv:2408.00181
-
EUDA: An Efficient Unsupervised Domain Adaptation via Self-Supervised Vision Transformer 31 Jul 2024 · 1 repository · arXiv:2407.21311
-
MimiQ: Low-Bit Data-Free Quantization of Vision Transformers with Encouraging Inter-Head Attention Similarity 29 Jul 2024 · 0 repositories · arXiv:2407.20021
-
Mixture of Nested Experts: Adaptive Processing of Visual Tokens 29 Jul 2024 · 1 repository · arXiv:2407.19985Syntology 8 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 1 honoured, 0 violated, 6 with no contract checked; 1 where Syntology's instrument failed) · 6 unverified (of 14 harvested samples)
-
Depth-Wise Convolutions in Vision Transformers for Efficient Training on Small Datasets 28 Jul 2024 · 1 repository · arXiv:2407.19394Syntology official (archive's flag): 13 ran · 13 ran (of which 0 constructed an object rather than computing a result; 11 with no instrument failure: 1 honoured, 2 violated, 8 with no contract checked; 2 where Syntology's instrument failed) · 3 unverified (of 16 harvested samples) · 16 pointer-only (licence)
-
Channel Boosted CNN-Transformer-based Multi-Level and Multi-Scale Nuclei Segmentation 27 Jul 2024 · 0 repositories · arXiv:2407.19186
-
Deep Companion Learning: Enhancing Generalization Through Historical Consistency 26 Jul 2024 · 0 repositories · arXiv:2407.18821
-
SHIC: Shape-Image Correspondences with no Keypoint Supervision 26 Jul 2024 · 0 repositories · arXiv:2407.18907
-
Skin Cancer Detection utilizing Deep Learning: Classification of Skin Lesion Images using a Vision Transformer 26 Jul 2024 · 0 repositories · arXiv:2407.18554
-
Case-Enhanced Vision Transformer: Improving Explanations of Image Similarity with a ViT-based Similarity Metric 24 Jul 2024 · 1 repository · arXiv:2407.16981
-
Graph Neural Networks: A suitable Alternative to MLPs in Latent 3D Medical Image Classification? 24 Jul 2024 · 1 repository · arXiv:2407.17219
-
Trans2Unet: Neural fusion for Nuclei Semantic Segmentation 24 Jul 2024 · 0 repositories · arXiv:2407.17181
-
Predicting the Best of N Visual Trackers 22 Jul 2024 · 1 repository · arXiv:2407.15707
-
Continual Distillation Learning: Knowledge Distillation in Prompt-based Continual Learning 18 Jul 2024 · 0 repositories · arXiv:2407.13911
-
LookupViT: Compressing visual information to a limited number of tokens 17 Jul 2024 · 0 repositories · arXiv:2407.12753
-
DiNO-Diffusion. Scaling Medical Diffusion via Self-Supervised Pre-Training 16 Jul 2024 · 0 repositories · arXiv:2407.11594
-
Probing the Efficacy of Federated Parameter-Efficient Fine-Tuning of Vision Transformers for Medical Image Classification 16 Jul 2024 · 0 repositories · arXiv:2407.11573
-
Siamese Transformer Networks for Few-shot Image Classification 16 Jul 2024 · 0 repositories · arXiv:2408.01427
-
Aligning Neuronal Coding of Dynamic Visual Scenes with Foundation Vision Models 15 Jul 2024 · 1 repository · arXiv:2407.10737