Methods › Computer Vision › Vision and Language Pre-Trained Models › CLIP › Papers where code ran, page 3
Contrastive Language-Image Pre-training
CLIP
Papers archive 2025-07-28
archive papers tagged: 3,094 · with a code link: 1,617 · where Syntology ran a sample: 649 (554 with a run with no instrument failure, 95 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (649 of 3,094 tagged: 554 with a run with no instrument failure, 95 where every run was a failure of Syntology's instrument)
Syntology We ran code from the paper's repository; we did not isolate this method inside it.
Page 3 of 7: papers 201 to 300 of the 649 tagged papers where Syntology ran at least one harvested sample (554 with a run with no instrument failure, 95 where every run was a failure of Syntology's instrument), newest first by the archive's date (ties by arXiv id). This is a filter on Syntology's record ordered by date only, not a ranking; a run is not a correctness claim. A paper missing from this list is not a recorded non-run: it may have no arXiv id, no harvested code, or only samples that have not run yet.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
On the test-time zero-shot generalization of vision-language models: Do we really need prompt learning? 3 May 2024 · 1 repository · arXiv:2405.02266Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
On Mechanistic Knowledge Localization in Text-to-Image Generative Models 2 May 2024 · 1 repository · arXiv:2405.01008Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
CLIPArTT: Adaptation of CLIP to New Domains at Test Time 1 May 2024 · 1 repository · arXiv:2405.00754Syntology official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 5 where Syntology's instrument failed) · 4 unverified (of 15 harvested samples) · 15 pointer-only (licence)
-
Weighted Point Cloud Embedding for Multimodal Contrastive Learning Toward Optimal Similarity Metric 30 Apr 2024 · 0 repositories · arXiv:2404.19228Syntology 5 ran (of which 4 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 7 harvested samples)
-
Revisiting the Adversarial Robustness of Vision Language Models: a Multimodal Perspective 30 Apr 2024 · 1 repository · arXiv:2404.19287Syntology official (archive's flag): 18 ran · 18 ran (of which 0 constructed an object rather than computing a result; 12 with no instrument failure: 1 honoured, 0 violated, 11 with no contract checked; 6 where Syntology's instrument failed) · 4 unverified (of 22 harvested samples) · 5 pointer-only (licence)
-
Modeling Caption Diversity in Contrastive Vision-Language Pretraining 30 Apr 2024 · 1 repository · arXiv:2405.00740Syntology official (archive's flag): 14 ran · 14 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 1 honoured, 0 violated, 9 with no contract checked; 4 where Syntology's instrument failed) · 8 unverified (of 22 harvested samples) · 22 pointer-only (licence)
-
Multi-Modal Proxy Learning Towards Personalized Visual Multiple Clustering 24 Apr 2024 · 1 repository · arXiv:2404.15655Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 3 where Syntology's instrument failed) · 3 unverified (of 7 harvested samples) · 7 pointer-only (licence)
-
MoVA: Adapting Mixture of Vision Experts to Multimodal Context 19 Apr 2024 · 1 repository · arXiv:2404.13046Syntology official (archive's flag): 8 ran · 9 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 1 violated, 3 with no contract checked; 5 where Syntology's instrument failed) · 2 unverified (of 11 harvested samples) · 3 pointer-only (licence)
-
The devil is in the object boundary: towards annotation-free instance segmentation using Foundation Models 18 Apr 2024 · 1 repository · arXiv:2404.11957Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
RankCLIP: Ranking-Consistent Language-Image Pretraining 15 Apr 2024 · 1 repository · arXiv:2404.09387Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 3 where Syntology's instrument failed) · 3 unverified (of 11 harvested samples) · 11 pointer-only (licence)
-
Leveraging Temporal Contextualization for Video Action Recognition 15 Apr 2024 · 2 repositories · arXiv:2404.09490Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Photo-Realistic Image Restoration in the Wild with Controlled Vision-Language Models 15 Apr 2024 · 2 repositories · arXiv:2404.09732Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 2 violated, 3 with no contract checked; 3 where Syntology's instrument failed) · 1 unverified (of 9 harvested samples) · 4 pointer-only (licence)
-
AMU-Tuning: Effective Logit Bias for CLIP-based Few-shot Learning 13 Apr 2024 · 1 repository · arXiv:2404.08958Syntology official (archive's flag): 4 ran · 4 ran (of which 2 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Pay Attention to Your Neighbours: Training-Free Open-Vocabulary Semantic Segmentation 12 Apr 2024 · 1 repository · arXiv:2404.08181Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 3 where Syntology's instrument failed) · 2 unverified (of 6 harvested samples) · 2 pointer-only (licence)
-
Two Effects, One Trigger: On the Modality Gap, Object Bias, and Information Imbalance in Contrastive Vision-Language Models 11 Apr 2024 · 1 repository · arXiv:2404.07983Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 4 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples)
-
View Selection for 3D Captioning via Diffusion Ranking 11 Apr 2024 · 2 repositories · arXiv:2404.07984Syntology 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Test-Time Adaptation with SaLIP: A Cascade of SAM and CLIP for Zero shot Medical Image Segmentation 9 Apr 2024 · 1 repository · arXiv:2404.06362Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 2 where Syntology's instrument failed) · 3 unverified (of 12 harvested samples) · 5 pointer-only (licence)
-
PURE: Turning Polysemantic Neurons Into Pure Features by Identifying Relevant Circuits 9 Apr 2024 · 1 repository · arXiv:2404.06453Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples)
-
To Cool or not to Cool? Temperature Network Meets Large Foundation Models via DRO 6 Apr 2024 · 1 repository · arXiv:2404.04575Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 3 where Syntology's instrument failed) · 3 unverified (of 8 harvested samples) · 8 pointer-only (licence)
-
Is CLIP the main roadblock for fine-grained open-world perception? 4 Apr 2024 · 2 repositories · arXiv:2404.03539Syntology official (archive's flag): 13 ran · 15 ran (of which 0 constructed an object rather than computing a result; 12 with no instrument failure: 0 honoured, 0 violated, 12 with no contract checked; 3 where Syntology's instrument failed) · 2 unverified (of 17 harvested samples) · 17 pointer-only (licence)
-
No "Zero-Shot" Without Exponential Data: Pretraining Concept Frequency Determines Multimodal Model Performance 4 Apr 2024 · 1 repository · arXiv:2404.04125Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 0 where Syntology's instrument failed) · 6 unverified (of 14 harvested samples)
-
LP++: A Surprisingly Strong Linear Probe for Few-Shot CLIP 2 Apr 2024 · 1 repository · arXiv:2404.02285Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 2 where Syntology's instrument failed) · 2 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
Do Vision-Language Models Understand Compound Nouns? 30 Mar 2024 · 1 repository · arXiv:2404.00419Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
CLAP4CLIP: Continual Learning with Probabilistic Finetuning for Vision-Language Models 28 Mar 2024 · 1 repository · arXiv:2403.19137Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
Beyond Embeddings: The Promise of Visual Table in Visual Reasoning 27 Mar 2024 · 1 repository · arXiv:2403.18252Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 1 honoured, 2 violated, 3 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 8 harvested samples) · 1 pointer-only (licence)
-
ImageNet-D: Benchmarking Neural Network Robustness on Diffusion Synthetic Object 27 Mar 2024 · 1 repository · arXiv:2403.18775Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 1 honoured, 0 violated, 2 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 6 harvested samples)
-
Dual Memory Networks: A Versatile Adaptation Approach for Vision-Language Models 26 Mar 2024 · 1 repository · arXiv:2403.17589Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 4 where Syntology's instrument failed) · 3 unverified (of 10 harvested samples) · 4 pointer-only (licence)
-
OmniVid: A Generative Framework for Universal Video Understanding 26 Mar 2024 · 1 repository · arXiv:2403.17935Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 1 where Syntology's instrument failed) · 4 unverified (of 11 harvested samples) · 11 pointer-only (licence)
-
Make-Your-Anchor: A Diffusion-based 2D Avatar Generation Framework 25 Mar 2024 · 1 repository · arXiv:2403.16510Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 2 where Syntology's instrument failed) · 2 unverified (of 8 harvested samples) · 8 pointer-only (licence)
-
DreamLIP: Language-Image Pre-training with Long Captions 25 Mar 2024 · 1 repository · arXiv:2403.17007Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 1 violated, 2 with no contract checked; 2 where Syntology's instrument failed) · 3 unverified (of 8 harvested samples) · 8 pointer-only (licence)
-
Continuous, Subject-Specific Attribute Control in T2I Models by Identifying Semantic Directions 25 Mar 2024 · 1 repository · arXiv:2403.17064Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples)
-
Task2Box: Box Embeddings for Modeling Asymmetric Task Relationships 25 Mar 2024 · 1 repository · arXiv:2403.17173Syntology official (archive's flag): 1 ran · 1 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified; the one sample that ran constructed an object rather than computing a result (of 1 harvested sample) · 1 pointer-only (licence)
-
Long-CLIP: Unlocking the Long-Text Capability of CLIP 22 Mar 2024 · 1 repository · arXiv:2403.15378Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 1 violated, 4 with no contract checked; 3 where Syntology's instrument failed) · 7 unverified (of 15 harvested samples) · 10 pointer-only (licence)
-
LLaVA-PruMerge: Adaptive Token Reduction for Efficient Large Multimodal Models 22 Mar 2024 · 1 repository · arXiv:2403.15388Syntology 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
C-TPT: Calibrated Test-Time Prompt Tuning for Vision-Language Models via Text Feature Dispersion 21 Mar 2024 · 1 repository · arXiv:2403.14119Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 6 where Syntology's instrument failed) · 2 unverified (of 10 harvested samples) · 5 pointer-only (licence)
-
OTSeg: Multi-prompt Sinkhorn Attention for Zero-Shot Semantic Segmentation 21 Mar 2024 · 1 repository · arXiv:2403.14183Syntology official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 1 honoured, 0 violated, 6 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 10 harvested samples) · 10 pointer-only (licence)
-
RAR: Retrieving And Ranking Augmented MLLMs for Visual Recognition 20 Mar 2024 · 2 repositories · arXiv:2403.13805Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 2 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 2 unverified (of 7 harvested samples) · 3 pointer-only (licence)
-
Adapting Visual-Language Models for Generalizable Anomaly Detection in Medical Images 19 Mar 2024 · 1 repository · arXiv:2403.12570Syntology official (archive's flag): 9 ran · 9 ran (of which 1 constructed an object rather than computing a result; 5 with no instrument failure: 1 honoured, 1 violated, 3 with no contract checked; 4 where Syntology's instrument failed) · 6 unverified (of 15 harvested samples) · 6 pointer-only (licence)
-
Boosting Continual Learning of Vision-Language Models via Mixture-of-Experts Adapters 18 Mar 2024 · 2 repositories · arXiv:2403.11549Syntology official (archive's flag): 2 ran · 2 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Arc2Face: A Foundation Model for ID-Consistent Human Faces 18 Mar 2024 · 3 repositories · arXiv:2403.11641Syntology community repositories only · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 4 harvested samples) · 3 pointer-only (licence)
-
Meta-Prompting for Automating Zero-shot Visual Recognition with LLMs 18 Mar 2024 · 1 repository · arXiv:2403.11755Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 3 where Syntology's instrument failed) · 2 unverified (of 7 harvested samples) · 3 pointer-only (licence)
-
Quality-Aware Image-Text Alignment for Real-World Image Quality Assessment 17 Mar 2024 · 1 repository · arXiv:2403.11176Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 2 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
MindEye2: Shared-Subject Models Enable fMRI-To-Image With 1 Hour of Data 17 Mar 2024 · 1 repository · arXiv:2403.11207Syntology official (archive's flag): 15 ran · 15 ran (of which 0 constructed an object rather than computing a result; 14 with no instrument failure: 0 honoured, 3 violated, 11 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 15 harvested samples) · 2 pointer-only (licence)
-
Unlocking the Multi-modal Potential of CLIP for Generalized Category Discovery 15 Mar 2024 · 1 repository · arXiv:2403.09974Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 5 where Syntology's instrument failed) · 7 unverified (of 14 harvested samples) · 3 pointer-only (licence)
-
Improving Medical Multi-modal Contrastive Learning with Expert Annotations 15 Mar 2024 · 1 repository · arXiv:2403.10153Syntology official (archive's flag): 5 ran · 5 ran (of which 2 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 2 where Syntology's instrument failed) · 7 unverified (of 12 harvested samples) · 12 pointer-only (licence)
-
CoLeCLIP: Open-Domain Continual Learning via Joint Task Prompt and Vocabulary Learning 15 Mar 2024 · 1 repository · arXiv:2403.10245Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples) · 1 pointer-only (licence)
-
Isotropic3D: Image-to-3D Generation Based on a Single CLIP Embedding 15 Mar 2024 · 1 repository · arXiv:2403.10395Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 3 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Calibrating Multi-modal Representations: A Pursuit of Group Robustness without Annotations 12 Mar 2024 · 1 repository · arXiv:2403.07241Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 3 harvested samples)
-
Beyond Text: Frozen Large Language Models in Visual Signal Comprehension 12 Mar 2024 · 1 repository · arXiv:2403.07874Syntology official (archive's flag): 17 ran · 17 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 1 honoured, 0 violated, 5 with no contract checked; 11 where Syntology's instrument failed) · 9 unverified (of 26 harvested samples) · 26 pointer-only (licence)
-
Toward Generalist Anomaly Detection via In-context Residual Learning with Few-shot Sample Prompts 11 Mar 2024 · 2 repositories · arXiv:2403.06495Syntology official (archive's flag): 16 ran · 16 ran (of which 0 constructed an object rather than computing a result; 14 with no instrument failure: 0 honoured, 1 violated, 13 with no contract checked; 2 where Syntology's instrument failed) · 5 unverified (of 21 harvested samples) · 16 pointer-only (licence)
-
Split to Merge: Unifying Separated Modalities for Unsupervised Domain Adaptation 11 Mar 2024 · 1 repository · arXiv:2403.06946Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 1 violated, 5 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 9 harvested samples) · 9 pointer-only (licence)
-
ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment 8 Mar 2024 · 2 repositories · arXiv:2403.05135Syntology 5 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 2 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples)
-
PEEB: Part-based Image Classifiers with an Explainable and Editable Language Bottleneck 8 Mar 2024 · 1 repository · arXiv:2403.05297Syntology official (archive's flag): 14 ran · 14 ran (of which 0 constructed an object rather than computing a result; 13 with no instrument failure: 0 honoured, 0 violated, 13 with no contract checked; 1 where Syntology's instrument failed) · 7 unverified (of 21 harvested samples)
-
Self-Adapting Large Visual-Language Models to Edge Devices across Visual Modalities 7 Mar 2024 · 1 repository · arXiv:2403.04908Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
FLAME Diffuser: Wildfire Image Synthesis using Mask Guided Diffusion 6 Mar 2024 · 1 repository · arXiv:2403.03463Syntology official (archive's flag): 13 ran · 13 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 6 where Syntology's instrument failed) · 3 unverified (of 16 harvested samples) · 16 pointer-only (licence)
-
MeaCap: Memory-Augmented Zero-shot Image Captioning 6 Mar 2024 · 1 repository · arXiv:2403.03715Syntology official (archive's flag): 6 ran · 6 ran (of which 1 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 2 where Syntology's instrument failed) · 4 unverified (of 10 harvested samples) · 10 pointer-only (licence)
-
What do we learn from inverting CLIP models? 5 Mar 2024 · 1 repository · arXiv:2403.02580Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
One Prompt Word is Enough to Boost Adversarial Robustness for Pre-trained Vision-Language Models 4 Mar 2024 · 1 repository · arXiv:2403.01849Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 6 harvested samples) · 1 pointer-only (licence)
-
G3DR: Generative 3D Reconstruction in ImageNet 1 Mar 2024 · 1 repository · arXiv:2403.00939Syntology official (archive's flag): 8 ran · 8 ran (of which 8 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 0 where Syntology's instrument failed) · 5 unverified; every one of the 8 samples that ran constructed an object rather than computing a result (of 13 harvested samples) · 13 pointer-only (licence)
-
Measuring Vision-Language STEM Skills of Neural Models 27 Feb 2024 · 1 repository · arXiv:2402.17205Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Fine-tuning CLIP Text Encoders with Two-step Paraphrasing 23 Feb 2024 · 0 repositories · arXiv:2402.15120Syntology 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Seeing is Believing: Mitigating Hallucination in Large Vision-Language Models via CLIP-Guided Decoding 23 Feb 2024 · 2 repositories · arXiv:2402.15300Syntology official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 11 with no instrument failure: 0 honoured, 0 violated, 11 with no contract checked; 0 where Syntology's instrument failed) · 6 unverified (of 17 harvested samples) · 17 pointer-only (licence)
-
Balanced Data Sampling for Language Model Training with Clustering 22 Feb 2024 · 1 repository · arXiv:2402.14526Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples)
-
CLIPping the Deception: Adapting Vision-Language Models for Universal Deepfake Detection 20 Feb 2024 · 1 repository · arXiv:2402.12927Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 8 harvested samples)
-
CounterCurate: Enhancing Physical and Semantic Visio-Linguistic Compositional Reasoning via Counterfactual Examples 20 Feb 2024 · 1 repository · arXiv:2402.13254Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 8 harvested samples) · 8 pointer-only (licence)
-
Robust CLIP: Unsupervised Adversarial Fine-Tuning of Vision Embeddings for Robust Large Vision-Language Models 19 Feb 2024 · 1 repository · arXiv:2402.12336Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 6 where Syntology's instrument failed) · 2 unverified (of 9 harvested samples)
-
ZeroG: Investigating Cross-dataset Zero-shot Transferability in Graphs 17 Feb 2024 · 1 repository · arXiv:2402.11235Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Interpreting CLIP with Sparse Linear Concept Embeddings (SpLiCE) 16 Feb 2024 · 1 repository · arXiv:2402.10376Syntology official (archive's flag): 7 ran · 7 ran (of which 1 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 10 harvested samples)
-
Open-Vocabulary Segmentation with Unpaired Mask-Text Supervision 14 Feb 2024 · 2 repositories · arXiv:2402.08960Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 7 harvested samples) · 7 pointer-only (licence)
-
ColorSwap: A Color and Word Order Dataset for Multimodal Evaluation 7 Feb 2024 · 1 repository · arXiv:2402.04492Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 2 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
OV-NeRF: Open-vocabulary Neural Radiance Fields with Vision and Language Foundation Models for 3D Semantic Understanding 7 Feb 2024 · 1 repository · arXiv:2402.04648Syntology official (archive's flag): 12 ran · 12 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 13 harvested samples) · 13 pointer-only (licence)
-
A Hard-to-Beat Baseline for Training-free CLIP-based Adaptation 6 Feb 2024 · 1 repository · arXiv:2402.04087Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 2 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters 6 Feb 2024 · 2 repositories · arXiv:2402.04252Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 6 harvested samples) · 4 pointer-only (licence)
-
Enhancing Compositional Generalization via Compositional Feature Alignment 5 Feb 2024 · 1 repository · arXiv:2402.02851Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples)
-
Variance Alignment Score: A Simple But Tough-to-Beat Data Selection Method for Multimodal Contrastive Learning 3 Feb 2024 · 2 repositories · arXiv:2402.02055Syntology 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 3 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
A Probabilistic Model Behind Self-Supervised Learning 2 Feb 2024 · 1 repository · arXiv:2402.01399Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 10 harvested samples) · 10 pointer-only (licence)
-
SynthCLIP: Are We Ready for a Fully Synthetic CLIP Training? 2 Feb 2024 · 1 repository · arXiv:2402.01832Syntology official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 3 where Syntology's instrument failed) · 1 unverified (of 12 harvested samples) · 12 pointer-only (licence)
-
Vision-LLMs Can Fool Themselves with Self-Generated Typographic Attacks 1 Feb 2024 · 1 repository · arXiv:2402.00626Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 1 where Syntology's instrument failed) · 4 unverified (of 8 harvested samples) · 8 pointer-only (licence)
-
Embracing Language Inclusivity and Diversity in CLIP through Continual Language Learning 30 Jan 2024 · 1 repository · arXiv:2401.17186Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples)
-
FreeStyle: Free Lunch for Text-guided Style Transfer using Diffusion Models 28 Jan 2024 · 1 repository · arXiv:2401.15636Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 1 where Syntology's instrument failed) · 3 unverified (of 7 harvested samples) · 7 pointer-only (licence)
-
Data-Free Generalized Zero-Shot Learning 28 Jan 2024 · 1 repository · arXiv:2401.15657Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 1 violated, 3 with no contract checked; 5 where Syntology's instrument failed) · 5 unverified (of 14 harvested samples) · 14 pointer-only (licence)
-
SciMMIR: Benchmarking Scientific Multi-modal Information Retrieval 24 Jan 2024 · 1 repository · arXiv:2401.13478Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 10 harvested samples) · 10 pointer-only (licence)
-
Prompting Large Vision-Language Models for Compositional Reasoning 20 Jan 2024 · 1 repository · arXiv:2401.11337Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
DGL: Dynamic Global-Local Prompt Tuning for Text-Video Retrieval 19 Jan 2024 · 2 repositories · arXiv:2401.10588Syntology official (archive's flag): 1 ran · 2 ran (of which 2 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified; every one of the 2 samples that ran constructed an object rather than computing a result (of 2 harvested samples) · 1 pointer-only (licence)
-
Supervised Fine-tuning in turn Improves Visual Foundation Models 18 Jan 2024 · 1 repository · arXiv:2401.10222Syntology official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 12 harvested samples)
-
Cross-modal Retrieval for Knowledge-based Visual Question Answering 11 Jan 2024 · 1 repository · arXiv:2401.05736Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 7 unverified (of 11 harvested samples) · 11 pointer-only (licence)
-
Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs 11 Jan 2024 · 1 repository · arXiv:2401.06209Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 1 violated, 2 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
Do Vision and Language Encoders Represent the World Similarly? 10 Jan 2024 · 1 repository · arXiv:2401.05224Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Towards Online Continuous Sign Language Recognition and Translation 10 Jan 2024 · 1 repository · arXiv:2401.05336Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Pre-trained Model Guided Fine-Tuning for Zero-Shot Adversarial Robustness 9 Jan 2024 · 1 repository · arXiv:2401.04350Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 10 harvested samples) · 10 pointer-only (licence)
-
Open-Vocabulary SAM: Segment and Recognize Twenty-thousand Classes Interactively 5 Jan 2024 · 1 repository · arXiv:2401.02955Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 1 violated, 4 with no contract checked; 2 where Syntology's instrument failed) · 4 unverified (of 11 harvested samples) · 11 pointer-only (licence)
-
Denoising Vision Transformers 5 Jan 2024 · 1 repository · arXiv:2401.02957Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 1 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 7 harvested samples) · 2 pointer-only (licence)
-
Latte: Latent Diffusion Transformer for Video Generation 5 Jan 2024 · 4 repositories · arXiv:2401.03048Syntology official (archive's flag): 12 ran · 12 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 1 honoured, 0 violated, 9 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 13 harvested samples)
-
Mining Fine-Grained Image-Text Alignment for Zero-Shot Captioning via Text-Only Training 4 Jan 2024 · 1 repository · arXiv:2401.02347Syntology official (archive's flag): 2 ran · 3 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 2 where Syntology's instrument failed) · 5 unverified (of 8 harvested samples) · 8 pointer-only (licence)
-
Learning to Prompt with Text Only Supervision for Vision-Language Models 4 Jan 2024 · 1 repository · arXiv:2401.02418Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 3 where Syntology's instrument failed) · 2 unverified (of 6 harvested samples) · 3 pointer-only (licence)
-
Improved Zero-Shot Classification by Adapting VLMs with Text Descriptions 4 Jan 2024 · 1 repository · arXiv:2401.02460Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
TinyGPT-V: Efficient Multimodal Large Language Model via Small Backbones 28 Dec 2023 · 2 repositories · arXiv:2312.16862Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 6 where Syntology's instrument failed) · 0 unverified (of 7 harvested samples)
-
Learning Vision from Models Rivals Learning Vision from Data 28 Dec 2023 · 2 repositories · arXiv:2312.17742Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 2 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 1 pointer-only (licence)
-
Forgery-aware Adaptive Transformer for Generalizable Synthetic Image Detection 27 Dec 2023 · 2 repositories · arXiv:2312.16649Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 5 where Syntology's instrument failed) · 1 unverified (of 7 harvested samples) · 4 pointer-only (licence)
-
HarmonyView: Harmonizing Consistency and Diversity in One-Image-to-3D 26 Dec 2023 · 1 repository · arXiv:2312.15980Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)