Methods › Computer Vision › Vision Transformers › Vision Transformer › Papers, page 7
Vision Transformer
Papers archive 2025-07-28
archive papers tagged: 2,144 · with a code link: 1,051 · where Syntology ran a sample: 328 (286 with a run with no instrument failure, 42 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (328 of 2,144 tagged: 286 with a run with no instrument failure, 42 where every run was a failure of Syntology's instrument)
Page 7 of 22: papers 601 to 700 of 2,144, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Anticipating Future Object Compositions without Forgetting 15 Jul 2024 · 0 repositories · arXiv:2407.10723
-
No Train, all Gain: Self-Supervised Gradients Improve Deep Frozen Representations 15 Jul 2024 · 1 repository · arXiv:2407.10964Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 2 where Syntology's instrument failed) · 5 unverified (of 12 harvested samples)
-
Unconstrained Open Vocabulary Image Classification: Zero-Shot Transfer from Text to Image via CLIP Inversion 15 Jul 2024 · 2 repositories · arXiv:2407.11211
-
Optimizing ROI Benefits Vehicle ReID in ITS 13 Jul 2024 · 0 repositories · arXiv:2407.09966
-
CXR-Agent: Vision-language models for chest X-ray interpretation with uncertainty aware radiology reporting 11 Jul 2024 · 0 repositories · arXiv:2407.08811
-
WildGaussians: 3D Gaussian Splatting in the Wild 11 Jul 2024 · 1 repository · arXiv:2407.08447Syntology official (archive's flag): 19 ran · 19 ran (of which 0 constructed an object rather than computing a result; 16 with no instrument failure: 2 honoured, 0 violated, 14 with no contract checked; 3 where Syntology's instrument failed) · 2 unverified (of 21 harvested samples) · 21 pointer-only (licence)
-
H-FCBFormer Hierarchical Fully Convolutional Branch Transformer for Occlusal Contact Segmentation with Articulating Paper 10 Jul 2024 · 1 repository · arXiv:2407.07604
-
Large Language Model-Augmented Auto-Delineation of Treatment Target Volume in Radiation Therapy 10 Jul 2024 · 0 repositories · arXiv:2407.07296
-
Swiss DINO: Efficient and Versatile Vision Framework for On-device Personal Object Search 10 Jul 2024 · 1 repository · arXiv:2407.07541
-
When to Accept Automated Predictions and When to Defer to Human Judgment? 10 Jul 2024 · 0 repositories · arXiv:2407.07821
-
Parameter-Efficient and Memory-Efficient Tuning for Vision Transformer: A Disentangled Approach 9 Jul 2024 · 1 repository · arXiv:2407.06964
-
Cross-domain Few-shot In-context Learning for Enhancing Traffic Sign Recognition 8 Jul 2024 · 0 repositories · arXiv:2407.05814
-
Multi-Label Plant Species Classification with Self-Supervised Vision Transformers 8 Jul 2024 · 1 repository · arXiv:2407.06298
-
Transfer Learning with Self-Supervised Vision Transformers for Snake Identification 8 Jul 2024 · 1 repository · arXiv:2407.06178
-
PRANCE: Joint Token-Optimization and Structural Channel-Pruning for Adaptive ViT Inference 6 Jul 2024 · 1 repository · arXiv:2407.05010
-
HCS-TNAS: Hybrid Constraint-driven Semi-supervised Transformer-NAS for Ultrasound Image Segmentation 5 Jul 2024 · 0 repositories · arXiv:2407.04203
-
Improving ensemble extreme precipitation forecasts using generative artificial intelligence 5 Jul 2024 · 0 repositories · arXiv:2407.04882
-
Multi-modal Masked Siamese Network Improves Chest X-Ray Representation Learning 5 Jul 2024 · 3 repositories · arXiv:2407.04449
-
Looking for Tiny Defects via Forward-Backward Feature Transfer 4 Jul 2024 · 0 repositories · arXiv:2407.04092
-
Self-supervised Vision Transformer are Scalable Generative Models for Domain Generalization 3 Jul 2024 · 1 repository · arXiv:2407.02900
-
Deep Learning Based Apparent Diffusion Coefficient Map Generation from Multi-parametric MR Images for Patients with Diffuse Gliomas 2 Jul 2024 · 0 repositories · arXiv:2407.02616
-
PoliFormer: Scaling On-Policy RL with Transformers Results in Masterful Navigators 28 Jun 2024 · 0 repositories · arXiv:2406.20083
-
Fibottention: Inceptive Visual Representation Learning with Diverse Attention Across Heads 27 Jun 2024 · 1 repository · arXiv:2406.19391
-
Segment Anything Model for automated image data annotation: empirical studies using text prompts from Grounding DINO 27 Jun 2024 · 0 repositories · arXiv:2406.19057
-
Human-Free Automated Prompting for Vision-Language Anomaly Detection: Prompt Optimization with Meta-guiding Prompt Scheme 26 Jun 2024 · 0 repositories · arXiv:2406.18197
-
Brain Tumor Classification using Vision Transformer with Selective Cross-Attention Mechanism and Feature Calibration 25 Jun 2024 · 0 repositories · arXiv:2406.17670
-
Semi-supervised classification of dental conditions in panoramic radiographs using large language model and instance segmentation: A real-world dataset evaluation 25 Jun 2024 · 0 repositories · arXiv:2406.17915
-
Task-Agnostic Federated Learning 25 Jun 2024 · 0 repositories · arXiv:2406.17235
-
Towards Optimal Trade-offs in Knowledge Distillation for CNNs and Vision Transformers at the Edge 25 Jun 2024 · 0 repositories · arXiv:2407.12808
-
Accelerating Phase Field Simulations Through a Hybrid Adaptive Fourier Neural Operator with U-Net Backbone 24 Jun 2024 · 0 repositories · arXiv:2406.17119
-
Diff3Dformer: Leveraging Slice Sequence Diffusion for Enhanced 3D CT Classification with Transformer Networks 24 Jun 2024 · 0 repositories · arXiv:2406.17173
-
Multi-Modal Vision Transformers for Crop Mapping from Satellite Image Time Series 24 Jun 2024 · 0 repositories · arXiv:2406.16513
-
Priorformer: A UGC-VQA Method with content and distortion priors 24 Jun 2024 · 0 repositories · arXiv:2406.16297
-
Breaking the Frame: Visual Place Recognition by Overlap Prediction 23 Jun 2024 · 1 repository · arXiv:2406.16204Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
SiT: Symmetry-Invariant Transformers for Generalisation in Reinforcement Learning 21 Jun 2024 · 1 repository · arXiv:2406.15025
-
SVFormer: A Direct Training Spiking Transformer for Efficient Video Action Recognition 21 Jun 2024 · 0 repositories · arXiv:2406.15034
-
Automatic Labels are as Effective as Manual Labels in Biomedical Images Classification with Deep Learning 20 Jun 2024 · 1 repository · arXiv:2406.14351
-
Enhanced Bank Check Security: Introducing a Novel Dataset and Transformer-Based Approach for Detection and Verification 20 Jun 2024 · 1 repository · arXiv:2406.14370
-
Guided Context Gating: Learning to leverage salient lesions in retinal fundus images 19 Jun 2024 · 0 repositories · arXiv:2406.13126
-
Liveness Detection in Computer Vision: Transformer-based Self-Supervised Learning for Face Anti-Spoofing 19 Jun 2024 · 0 repositories · arXiv:2406.13860
-
MixDiff: Mixing Natural and Synthetic Images for Robust Self-Supervised Representations 18 Jun 2024 · 1 repository · arXiv:2406.12368
-
PCIE_EgoHandPose Solution for EgoExo4D Hand Pose Challenge 18 Jun 2024 · 1 repository · arXiv:2406.12219
-
Inpainting the Gaps: A Novel Framework for Evaluating Explanation Methods in Vision Transformers 17 Jun 2024 · 0 repositories · arXiv:2406.11534
-
Object Detection using Oriented Window Learning Vi-sion Transformer: Roadway Assets Recognition 15 Jun 2024 · 0 repositories · arXiv:2406.10712
-
Self-Supervised Vision Transformer for Enhanced Virtual Clothes Try-On 15 Jun 2024 · 0 repositories · arXiv:2406.10539
-
When Will Gradient Regularization Be Harmful? 14 Jun 2024 · 1 repository · arXiv:2406.09723Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
An Image is Worth More Than 16x16 Patches: Exploring Transformers on Individual Pixels 13 Jun 2024 · 0 repositories · arXiv:2406.09415
-
Fusion of regional and sparse attention in Vision Transformers 13 Jun 2024 · 0 repositories · arXiv:2406.08859
-
MGRQ: Post-Training Quantization For Vision Transformer With Mixed Granularity Reconstruction 13 Jun 2024 · 0 repositories · arXiv:2406.09229
-
Parameter-Efficient Active Learning for Foundational models 13 Jun 2024 · 0 repositories · arXiv:2406.09296
-
AdaNCA: Neural Cellular Automata As Adaptors For More Robust Vision Transformer 12 Jun 2024 · 0 repositories · arXiv:2406.08298
-
ConceptHash: Interpretable Fine-Grained Hashing via Concept Discovery 12 Jun 2024 · 1 repository · arXiv:2406.08457
-
ICE-G: Image Conditional Editing of 3D Gaussian Splats 12 Jun 2024 · 0 repositories · arXiv:2406.08488
-
GridPE: Unifying Positional Encoding in Transformers with a Grid Cell-Inspired Framework 11 Jun 2024 · 0 repositories · arXiv:2406.07049
-
UVIS: Unsupervised Video Instance Segmentation 11 Jun 2024 · 0 repositories · arXiv:2406.06908
-
A Comparative Survey of Vision Transformers for Feature Extraction in Texture Analysis 10 Jun 2024 · 0 repositories · arXiv:2406.06136
-
GCtx-UNet: Efficient Network for Medical Image Segmentation 9 Jun 2024 · 1 repository · arXiv:2406.05891
-
1st Place Winner of the 2024 Pixel-level Video Understanding in the Wild (CVPR'24 PVUW) Challenge in Video Panoptic Segmentation and Best Long Video Consistency of Video Semantic Segmentation 8 Jun 2024 · 0 repositories · arXiv:2406.05352
-
U-Net Ensemble for Enhanced Semantic Segmentation in Remote Sensing Imagery 8 Jun 2024 · 0 repositories
-
REP: Resource-Efficient Prompting for Rehearsal-Free Continual Learning 7 Jun 2024 · 0 repositories · arXiv:2406.04772
-
DeepStack: Deeply Stacking Visual Tokens is Surprisingly Simple and Effective for LMMs 6 Jun 2024 · 0 repositories · arXiv:2406.04334
-
ReDistill: Residual Encoded Distillation for Peak Memory Reduction 6 Jun 2024 · 0 repositories · arXiv:2406.03744
-
Learning Visual Prompts for Guiding the Attention of Vision Transformers 5 Jun 2024 · 0 repositories · arXiv:2406.03303
-
Stable-Pose: Leveraging Transformers for Pose-Guided Text-to-Image Generation 4 Jun 2024 · 1 repository · arXiv:2406.02485Syntology official (archive's flag): 12 ran · 12 ran (of which 0 constructed an object rather than computing a result; 11 with no instrument failure: 1 honoured, 4 violated, 6 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 13 harvested samples) · 13 pointer-only (licence)
-
The Deep Latent Space Particle Filter for Real-Time Data Assimilation with Uncertainty Quantification 4 Jun 2024 · 1 repository · arXiv:2406.02204
-
Decomposing and Interpreting Image Representations via Text in ViTs Beyond CLIP 3 Jun 2024 · 1 repository · arXiv:2406.01583Syntology official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 1 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 13 harvested samples) · 13 pointer-only (licence)
-
ELSA: Evaluating Localization of Social Activities in Urban Streets using Open-Vocabulary Detection 3 Jun 2024 · 0 repositories · arXiv:2406.01551
-
Eating Smart: Advancing Health Informatics with the Grounding DINO based Dietary Assistant App 2 Jun 2024 · 0 repositories · arXiv:2406.00848
-
Kolmogorov-Arnold Network for Satellite Image Classification in Remote Sensing 2 Jun 2024 · 1 repository · arXiv:2406.00600
-
MGI: Multimodal Contrastive pre-training of Genomic and Medical Imaging 2 Jun 2024 · 0 repositories · arXiv:2406.00631
-
A Deep Learning Model for Coronary Artery Segmentation and Quantitative Stenosis Detection in Angiographic Images 1 Jun 2024 · 1 repository · arXiv:2406.00492
-
You Only Need Less Attention at Each Stage in Vision Transformers 1 Jun 2024 · 0 repositories · arXiv:2406.00427
-
MVAD: A Multiple Visual Artifact Detector for Video Streaming 31 May 2024 · 0 repositories · arXiv:2406.00212
-
Ovis: Structural Embedding Alignment for Multimodal Large Language Model 31 May 2024 · 2 repositories · arXiv:2405.20797
-
Use of a Multiscale Vision Transformer to predict Nursing Activities Score from Low Resolution Thermal Videos in an Intensive Care Unit 30 May 2024 · 0 repositories · arXiv:2406.04364
-
Enhancing Vision-Language Model with Unmasked Token Alignment 29 May 2024 · 1 repository · arXiv:2405.19009
-
FDQN: A Flexible Deep Q-Network Framework for Game Automation 29 May 2024 · 1 repository · arXiv:2405.18761
-
MDS-ViTNet: Improving saliency prediction for Eye-Tracking with Vision Transformer 29 May 2024 · 1 repository · arXiv:2405.19501
-
Adapting Pre-Trained Vision Models for Novel Instance Detection and Segmentation 28 May 2024 · 1 repository · arXiv:2405.17859Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 7 harvested samples)
-
Efficient Time Series Processing for Transformers and State-Space Models through Token Merging 28 May 2024 · 0 repositories · arXiv:2405.17951
-
Near-Infrared and Low-Rank Adaptation of Vision Transformers in Remote Sensing 28 May 2024 · 0 repositories · arXiv:2405.17901
-
Visual Anchors Are Strong Information Aggregators For Multimodal Large Language Model 28 May 2024 · 1 repository · arXiv:2405.17815
-
Visualizing the loss landscape of Self-supervised Vision Transformer 28 May 2024 · 0 repositories · arXiv:2405.18042
-
Wavelet-Based Image Tokenizer for Vision Transformers 28 May 2024 · 0 repositories · arXiv:2405.18616
-
How Do the Architecture and Optimizer Affect Representation Learning? On the Training Dynamics of Representations in Deep Neural Networks 27 May 2024 · 0 repositories · arXiv:2405.17377
-
SA-GS: Semantic-Aware Gaussian Splatting for Large Scene Reconstruction with Geometry Constrain 27 May 2024 · 0 repositories · arXiv:2405.16923
-
Supervised Batch Normalization 27 May 2024 · 0 repositories · arXiv:2405.17027
-
ConvLLaVA: Hierarchical Backbones as Visual Encoder for Large Multimodal Models 24 May 2024 · 1 repository · arXiv:2405.15738Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 3 where Syntology's instrument failed) · 2 unverified (of 10 harvested samples) · 1 pointer-only (licence)
-
Recasting Generic Pretrained Vision Transformers As Object-Centric Scene Encoders For Manipulation Policies 24 May 2024 · 0 repositories · arXiv:2405.15916
-
Steerable Transformers 24 May 2024 · 0 repositories · arXiv:2405.15932
-
Designing A Sustainable Marine Debris Clean-up Framework without Human Labels 23 May 2024 · 1 repository · arXiv:2405.14815
-
Magnetic Resonance Image Processing Transformer for General Accelerated Image Reconstruction 23 May 2024 · 0 repositories · arXiv:2405.15098
-
PrivCirNet: Efficient Private Inference via Block Circulant Transformation 23 May 2024 · 1 repository · arXiv:2405.14569Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 2 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples)
-
Sparse-Tuning: Adapting Vision Transformers with Efficient Fine-tuning and Inference 23 May 2024 · 1 repository · arXiv:2405.14700Syntology official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 1 honoured, 0 violated, 8 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 11 harvested samples) · 11 pointer-only (licence)
-
CViT: Continuous Vision Transformer for Operator Learning 22 May 2024 · 2 repositories · arXiv:2405.13998Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 4 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples)
-
LookHere: Vision Transformers with Directed Attention Generalize and Extrapolate 22 May 2024 · 1 repository · arXiv:2405.13985Syntology official (archive's flag): 8 ran · 8 ran (of which 5 constructed an object rather than computing a result; 7 with no instrument failure: 2 honoured, 0 violated, 5 with no contract checked; 1 where Syntology's instrument failed) · 4 unverified (of 12 harvested samples)
-
Semantic Equitable Clustering: A Simple and Effective Strategy for Clustering Vision Tokens 22 May 2024 · 0 repositories · arXiv:2405.13337
-
Text Prompting for Multi-Concept Video Customization by Autoregressive Generation 22 May 2024 · 0 repositories · arXiv:2405.13951
-
BIMM: Brain Inspired Masked Modeling for Video Representation Learning 21 May 2024 · 1 repository · arXiv:2405.12757
-
Is Dataset Quality Still a Concern in Diagnosis Using Large Foundation Model? 21 May 2024 · 0 repositories · arXiv:2405.12584