Methods › Computer Vision › Vision Transformers › Vision Transformer › Papers, page 8
Vision Transformer
Papers archive 2025-07-28
archive papers tagged: 2,144 · with a code link: 1,051 · where Syntology ran a sample: 328 (286 with a run with no instrument failure, 42 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (328 of 2,144 tagged: 286 with a run with no instrument failure, 42 where every run was a failure of Syntology's instrument)
Page 8 of 22: papers 701 to 800 of 2,144, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Global-local Fourier Neural Operator for Accelerating Coronal Magnetic Field Model 21 May 2024 · 1 repository · arXiv:2405.12754Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample)
-
A Method on Searching Better Activation Functions 19 May 2024 · 0 repositories · arXiv:2405.12954
-
Track Anything Rapter(TAR) 19 May 2024 · 1 repository · arXiv:2405.11655
-
Towards SAR Automatic Target Recognition MultiCategory SAR Image Classification Based on Light Weight Vision Transformer 18 May 2024 · 0 repositories · arXiv:2407.06128
-
DINO as a von Mises-Fisher mixture model 17 May 2024 · 0 repositories · arXiv:2405.10939
-
Enhancing the analysis of murine neonatal ultrasonic vocalizations: Development, evaluation, and application of different mathematical models 17 May 2024 · 1 repository · arXiv:2405.12957
-
Grounding DINO 1.5: Advance the "Edge" of Open-Set Object Detection 16 May 2024 · 3 repositories · arXiv:2405.10300Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples)
-
A Comprehensive Evaluation of Histopathology Foundation Models for Ovarian Cancer Subtype Classification 16 May 2024 · 1 repository · arXiv:2405.09990
-
Quantum Vision Transformers for Quark-Gluon Classification 16 May 2024 · 1 repository · arXiv:2405.10284Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Vision Transformers for End-to-End Vision-Based Quadrotor Obstacle Avoidance 16 May 2024 · 0 repositories · arXiv:2405.10391
-
Perception- and Fidelity-aware Reduced-Reference Super-Resolution Image Quality Assessment 15 May 2024 · 0 repositories · arXiv:2405.09472
-
A Timely Survey on Vision Transformer for Deepfake Detection 14 May 2024 · 0 repositories · arXiv:2405.08463
-
Abnormal Respiratory Sound Identification Using Audio-Spectrogram Vision Transformer 14 May 2024 · 0 repositories · arXiv:2405.08342
-
Rethinking Scanning Strategies with Vision Mamba in Semantic Segmentation of Remote Sensing Imagery: An Experimental Study 14 May 2024 · 0 repositories · arXiv:2405.08493
-
NutritionVerse-Direct: Exploring Deep Neural Networks for Multitask Nutrition Prediction from Food Images 13 May 2024 · 0 repositories · arXiv:2405.07814
-
BoQ: A Place is Worth a Bag of Learnable Queries 12 May 2024 · 1 repository · arXiv:2405.07364Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
QMViT: A Mushroom is worth 16x16 Words 11 May 2024 · 0 repositories · arXiv:2407.04708
-
Dual-Task Vision Transformer for Rapid and Accurate Intracerebral Hemorrhage CT Image Classification 10 May 2024 · 0 repositories · arXiv:2405.06814
-
An Advanced Features Extraction Module for Remote Sensing Image Super-Resolution 7 May 2024 · 0 repositories · arXiv:2405.04595
-
Structured Click Control in Transformer-based Interactive Segmentation 7 May 2024 · 1 repository · arXiv:2405.04009Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 1 honoured, 0 violated, 5 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 9 harvested samples) · 9 pointer-only (licence)
-
Class-relevant Patch Embedding Selection for Few-Shot Image Classification 6 May 2024 · 0 repositories · arXiv:2405.03722
-
Swin transformers are robust to distribution and concept drift in endoscopy-based longitudinal rectal cancer assessment 6 May 2024 · 0 repositories · arXiv:2405.03762
-
Intra-task Mutual Attention based Vision Transformer for Few-Shot Learning 6 May 2024 · 0 repositories · arXiv:2405.03109
-
Boosting 3D Neuron Segmentation with 2D Vision Transformer Pre-trained on Natural Images 4 May 2024 · 0 repositories · arXiv:2405.02686
-
An Attention Based Pipeline for Identifying Pre-Cancer Lesions in Head and Neck Clinical Images 3 May 2024 · 1 repository · arXiv:2405.01937
-
Multi-method Integration with Confidence-based Weighting for Zero-shot Image Classification 3 May 2024 · 0 repositories · arXiv:2405.02155
-
Torch2Chip: An End-to-end Customizable Deep Neural Network Compression and Deployment Toolkit for Prototype Hardware Accelerator Design 2 May 2024 · 1 repository · arXiv:2405.01775
-
Brighteye: Glaucoma Screening with Color Fundus Photographs based on Vision Transformer 1 May 2024 · 1 repository · arXiv:2405.00857
-
Exploring Self-Supervised Vision Transformers for Deepfake Detection: A Comparative Analysis 1 May 2024 · 1 repository · arXiv:2405.00355
-
LOTUS: Improving Transformer Efficiency with Sparsity Pruning and Data Lottery Tickets 1 May 2024 · 0 repositories · arXiv:2405.00906
-
Automatic Cardiac Pathology Recognition in Echocardiography Images Using Higher Order Dynamic Mode Decomposition and a Vision Transformer for Small Datasets 30 Apr 2024 · 0 repositories · arXiv:2404.19579
-
CLIP-Mamba: CLIP Pretrained Mamba Models with OOD and Hessian Evaluation 30 Apr 2024 · 1 repository · arXiv:2404.19394
-
Masked Multi-Query Slot Attention for Unsupervised Object Discovery 30 Apr 2024 · 1 repository · arXiv:2404.19654
-
Neuro-Vision to Language: Enhancing Brain Recording-based Visual Reconstruction and Language Interaction 30 Apr 2024 · 0 repositories · arXiv:2404.19438
-
Seeing Through the Clouds: Cloud Gap Imputation with Prithvi Foundation Model 30 Apr 2024 · 1 repository · arXiv:2404.19609
-
Harmonic Machine Learning Models are Robust 29 Apr 2024 · 0 repositories · arXiv:2404.18825
-
Fashion Recommendation: Outfit Compatibility using GNN 28 Apr 2024 · 1 repository · arXiv:2404.18040
-
MultiMAE-DER: Multimodal Masked Autoencoder for Dynamic Emotion Recognition 28 Apr 2024 · 1 repository · arXiv:2404.18327
-
CLFT: Camera-LiDAR Fusion Transformer for Semantic Segmentation in Autonomous Driving 27 Apr 2024 · 2 repositories · arXiv:2404.17793
-
Binarizing Documents by Leveraging both Space and Frequency 26 Apr 2024 · 1 repository · arXiv:2404.17243
-
Parameter Efficient Fine-tuning of Self-supervised ViTs without Catastrophic Forgetting 26 Apr 2024 · 1 repository · arXiv:2404.17245
-
Image Quality Assessment With Compressed Sampling 26 Apr 2024 · 0 repositories · arXiv:2404.17170
-
SAGHOG: Self-Supervised Autoencoder for Generating HOG Features for Writer Retrieval 26 Apr 2024 · 1 repository · arXiv:2404.17221
-
UniRGB-IR: A Unified Framework for RGB-Infrared Semantic Tasks via Adapter Tuning 26 Apr 2024 · 1 repository · arXiv:2404.17360
-
Boosting Unsupervised Semantic Segmentation with Principal Mask Proposals 25 Apr 2024 · 1 repository · arXiv:2404.16818
-
TinyChart: Efficient Chart Understanding with Visual Token Merging and Program-of-Thoughts Learning 25 Apr 2024 · 1 repository · arXiv:2404.16635
-
MiM: Mask in Mask Self-Supervised Pre-Training for 3D Medical Image Analysis 24 Apr 2024 · 0 repositories · arXiv:2404.15580
-
Rethinking model prototyping through the MedMNIST+ dataset collection 24 Apr 2024 · 1 repository · arXiv:2404.15786
-
SPARO: Selective Attention for Robust and Compositional Transformer Encodings for Vision 24 Apr 2024 · 1 repository · arXiv:2404.15721
-
Vision Transformer-based Adversarial Domain Adaptation 24 Apr 2024 · 1 repository · arXiv:2404.15817
-
ORBIT: Oak Ridge Base Foundation Model for Earth System Predictability 23 Apr 2024 · 0 repositories · arXiv:2404.14712
-
ThermoPore: Predicting Part Porosity Based on Thermal Images Using Deep Learning 23 Apr 2024 · 0 repositories · arXiv:2404.16882
-
1st Place Solution to the 1st SkatingVerse Challenge 22 Apr 2024 · 0 repositories · arXiv:2404.14032
-
Cross-Task Multi-Branch Vision Transformer for Facial Expression and Mask Wearing Classification 22 Apr 2024 · 0 repositories · arXiv:2404.14606
-
FiLo: Zero-Shot Anomaly Detection by Fine-Grained Description and High-Quality Localization 21 Apr 2024 · 1 repository · arXiv:2404.13671Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 1 violated, 4 with no contract checked; 2 where Syntology's instrument failed) · 5 unverified (of 12 harvested samples) · 7 pointer-only (licence)
-
LMFNet: An Efficient Multimodal Fusion Approach for Semantic Segmentation in High-Resolution Remote Sensing 21 Apr 2024 · 0 repositories · arXiv:2404.13659
-
Masked Latent Transformer with the Random Masking Ratio to Advance the Diagnosis of Dental Fluorosis 21 Apr 2024 · 1 repository · arXiv:2404.13564
-
Vim4Path: Self-Supervised Vision Mamba for Histopathology Images 20 Apr 2024 · 1 repository · arXiv:2404.13222Syntology official (archive's flag): 1 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 1 pointer-only (licence)
-
Towards Robust Ferrous Scrap Material Classification with Deep Learning and Conformal Prediction 19 Apr 2024 · 0 repositories · arXiv:2404.13002
-
The devil is in the object boundary: towards annotation-free instance segmentation using Foundation Models 18 Apr 2024 · 1 repository · arXiv:2404.11957Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
JointViT: Modeling Oxygen Saturation Levels with Joint Supervision on Long-Tailed OCTA 17 Apr 2024 · 1 repository · arXiv:2404.11525
-
Pretraining Billion-scale Geospatial Foundational Models on Frontier 17 Apr 2024 · 0 repositories · arXiv:2404.11706
-
Supervised Contrastive Vision Transformer for Breast Histopathological Image Classification 17 Apr 2024 · 0 repositories · arXiv:2404.11052
-
Gasformer: A Transformer-based Architecture for Segmenting Methane Emissions from Livestock in Optical Gas Imaging 16 Apr 2024 · 1 repository · arXiv:2404.10841
-
Arena: A Patch-of-Interest ViT Inference Acceleration System for Edge-Assisted Video Analytics 14 Apr 2024 · 0 repositories · arXiv:2404.09245
-
A Novel Vision Transformer based Load Profile Analysis using Load Images as Inputs 12 Apr 2024 · 0 repositories · arXiv:2404.08175
-
IFViT: Interpretable Fixed-Length Representation for Fingerprint Matching via Vision Transformer 12 Apr 2024 · 0 repositories · arXiv:2404.08237
-
Pay Attention to Your Neighbours: Training-Free Open-Vocabulary Semantic Segmentation 12 Apr 2024 · 1 repository · arXiv:2404.08181Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 3 where Syntology's instrument failed) · 2 unverified (of 6 harvested samples) · 2 pointer-only (licence)
-
Single-image driven 3d viewpoint training data augmentation for effective wine label recognition 12 Apr 2024 · 0 repositories · arXiv:2404.08820
-
Progressive Semantic-Guided Vision Transformer for Zero-Shot Learning 11 Apr 2024 · 1 repository · arXiv:2404.07713Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 8 harvested samples) · 8 pointer-only (licence)
-
InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD 9 Apr 2024 · 2 repositories · arXiv:2404.06512
-
WebCode2M: A Real-World Dataset for Code Generation from Webpage Designs 9 Apr 2024 · 0 repositories · arXiv:2404.06369
-
HSViT: Horizontally Scalable Vision Transformer 8 Apr 2024 · 1 repository · arXiv:2404.05196
-
GvT: A Graph-based Vision Transformer with Talking-Heads Utilizing Sparsity, Trained from Scratch on Small Datasets 7 Apr 2024 · 0 repositories · arXiv:2404.04924
-
Hyperbolic Learning with Synthetic Captions for Open-World Detection 7 Apr 2024 · 0 repositories · arXiv:2404.05016
-
VMambaMorph: a Multi-Modality Deformable Image Registration Framework based on Visual State Space Model with Cross-Scan Module 7 Apr 2024 · 1 repository · arXiv:2404.05105
-
Cluster-based Video Summarization with Temporal Context Awareness 6 Apr 2024 · 1 repository · arXiv:2404.04511
-
Learning Correlation Structures for Vision Transformers 5 Apr 2024 · 0 repositories · arXiv:2404.03924
-
OpenNeRF: Open Set 3D Neural Scene Segmentation with Pixel-Wise Features and Rendered Novel Views 4 Apr 2024 · 0 repositories · arXiv:2404.03650
-
Minimize Quantization Output Error with Bias Compensation 2 Apr 2024 · 1 repository · arXiv:2404.01892
-
PREGO: online mistake detection in PRocedural EGOcentric videos 2 Apr 2024 · 1 repository · arXiv:2404.01933Syntology official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 1 violated, 7 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 11 harvested samples) · 3 pointer-only (licence)
-
Samba: Semantic Segmentation of Remotely Sensed Images with State Space Model 2 Apr 2024 · 1 repository · arXiv:2404.01705Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 8 harvested samples) · 7 pointer-only (licence)
-
Can Biases in ImageNet Models Explain Generalization? 1 Apr 2024 · 1 repository · arXiv:2404.01509Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 9 harvested samples) · 9 pointer-only (licence)
-
Flare-Free Vision: Empowering Uformer with Depth Insights 1 Apr 2024 · 1 repository
-
Open-Vocabulary Object Detectors: Robustness Challenges under Distribution Shifts 1 Apr 2024 · 0 repositories · arXiv:2405.14874
-
On the Faithfulness of Vision Transformer Explanations 1 Apr 2024 · 0 repositories · arXiv:2404.01415
-
Structured Initialization for Attention in Vision Transformers 1 Apr 2024 · 1 repository · arXiv:2404.01139
-
Vision-language models for decoding provider attention during neonatal resuscitation 1 Apr 2024 · 0 repositories · arXiv:2404.01207
-
AgileFormer: Spatially Agile Transformer UNet for Medical Image Segmentation 29 Mar 2024 · 1 repository · arXiv:2404.00122
-
Enhancing Efficiency in Vision Transformer Networks: Design Techniques and Insights 28 Mar 2024 · 0 repositories · arXiv:2403.19882
-
Patch Spatio-Temporal Relation Prediction for Video Anomaly Detection 28 Mar 2024 · 0 repositories · arXiv:2403.19111
-
Siamese Vision Transformers are Scalable Audio-visual Learners 28 Mar 2024 · 1 repository · arXiv:2403.19638Syntology official (archive's flag): 15 ran · 15 ran (of which 3 constructed an object rather than computing a result; 15 with no instrument failure: 0 honoured, 0 violated, 15 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 19 harvested samples) · 19 pointer-only (licence)
-
ECoDepth: Effective Conditioning of Diffusion Models for Monocular Depth Estimation 27 Mar 2024 · 1 repository · arXiv:2403.18807Syntology official (archive's flag): 11 ran · 11 ran (of which 1 constructed an object rather than computing a result; 6 with no instrument failure: 2 honoured, 2 violated, 2 with no contract checked; 5 where Syntology's instrument failed) · 4 unverified (of 15 harvested samples) · 15 pointer-only (licence)
-
Illicit object detection in X-ray images using Vision Transformers 27 Mar 2024 · 0 repositories · arXiv:2403.19043
-
Lift3D: Zero-Shot Lifting of Any 2D Vision Model to 3D 27 Mar 2024 · 0 repositories · arXiv:2403.18922
-
ViTAR: Vision Transformer with Any Resolution 27 Mar 2024 · 0 repositories · arXiv:2403.18361
-
Accuracy enhancement method for speech emotion recognition from spectrogram using temporal frequency correlation and positional information learning through knowledge transfer 26 Mar 2024 · 1 repository · arXiv:2403.17327
-
Evaluating the Efficacy of Prompt-Engineered Large Multimodal Models Versus Fine-Tuned Vision Transformers in Image-Based Security Applications 26 Mar 2024 · 0 repositories · arXiv:2403.17787
-
3D-EffiViTCaps: 3D Efficient Vision Transformer with Capsule for Medical Image Segmentation 25 Mar 2024 · 1 repository · arXiv:2403.16350
-
DTF-AT: Decoupled Time-Frequency Audio Transformer for Event Classification 24 Mar 2024 · 1 repository