Methods › Computer Vision › Vision and Language Pre-Trained Models › CLIP › Papers, page 15
Contrastive Language-Image Pre-training
CLIP
Papers archive 2025-07-28
archive papers tagged: 3,094 · with a code link: 1,617 · where Syntology ran a sample: 649 (554 with a run with no instrument failure, 95 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (649 of 3,094 tagged: 554 with a run with no instrument failure, 95 where every run was a failure of Syntology's instrument)
Page 15 of 31: papers 1,401 to 1,500 of 3,094, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Zero-Shot Distillation for Image Encoders: How to Make Effective Use of Synthetic Data 25 Apr 2024 · 0 repositories · arXiv:2404.16637
-
FairDeDup: Detecting and Mitigating Vision-Language Fairness Disparities in Semantic Dataset Deduplication 24 Apr 2024 · 0 repositories · arXiv:2404.16123
-
Mammo-CLIP: Leveraging Contrastive Language-Image Pre-training (CLIP) for Enhanced Breast Cancer Diagnosis with Multi-view Mammography 24 Apr 2024 · 0 repositories · arXiv:2404.15946
-
MoDE: CLIP Data Experts via Clustering 24 Apr 2024 · 1 repository · arXiv:2404.16030
-
Multi-Modal Proxy Learning Towards Personalized Visual Multiple Clustering 24 Apr 2024 · 1 repository · arXiv:2404.15655Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 3 where Syntology's instrument failed) · 3 unverified (of 7 harvested samples) · 7 pointer-only (licence)
-
Seeing Beyond Classes: Zero-Shot Grounded Situation Recognition via Language Explainer 24 Apr 2024 · 0 repositories · arXiv:2404.15785
-
SPARO: Selective Attention for Robust and Compositional Transformer Encodings for Vision 24 Apr 2024 · 1 repository · arXiv:2404.15721
-
Adaptive Prompt Learning with Negative Textual Semantics and Uncertainty Modeling for Universal Multi-Source Domain Adaptation 23 Apr 2024 · 0 repositories · arXiv:2404.14696
-
CT-GLIP: 3D Grounded Language-Image Pretraining with CT Scans and Radiology Reports for Full-Body Scenarios 23 Apr 2024 · 0 repositories · arXiv:2404.15272
-
Multi-Modal Prompt Learning on Blind Image Quality Assessment 23 Apr 2024 · 1 repository · arXiv:2404.14949
-
CLIP-GS: CLIP-Informed Gaussian Splatting for Real-time and View-consistent 3D Semantic Understanding 22 Apr 2024 · 1 repository · arXiv:2404.14249
-
WangLab at MEDIQA-M3G 2024: Multimodal Medical Answer Generation using Large Language Models 22 Apr 2024 · 0 repositories · arXiv:2404.14567
-
Hyper-SD: Trajectory Segmented Consistency Model for Efficient Image Synthesis 21 Apr 2024 · 0 repositories · arXiv:2404.13686
-
Iteratively Prompting Multimodal LLMs to Reproduce Natural and AI-Generated Images 21 Apr 2024 · 0 repositories · arXiv:2404.13784
-
Object-Attribute Binding in Text-to-Image Generation: Evaluation and Control 21 Apr 2024 · 0 repositories · arXiv:2404.13766
-
PCQA: A Strong Baseline for AIGC Quality Assessment Based on Prompt Condition 20 Apr 2024 · 0 repositories · arXiv:2404.13299
-
Data Alignment for Zero-Shot Concept Generation in Dermatology AI 19 Apr 2024 · 0 repositories · arXiv:2404.13043
-
ECOR: Explainable CLIP for Object Recognition 19 Apr 2024 · 0 repositories · arXiv:2404.12839
-
Exploring Interactive Semantic Alignment for Efficient HOI Detection with Vision-language Model 19 Apr 2024 · 0 repositories · arXiv:2404.12678
-
MoVA: Adapting Mixture of Vision Experts to Multimodal Context 19 Apr 2024 · 1 repository · arXiv:2404.13046Syntology official (archive's flag): 8 ran · 9 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 1 violated, 3 with no contract checked; 5 where Syntology's instrument failed) · 2 unverified (of 11 harvested samples) · 3 pointer-only (licence)
-
Robust CLIP-Based Detector for Exposing Diffusion Model-Generated Images 19 Apr 2024 · 1 repository · arXiv:2404.12908
-
Unified Scene Representation and Reconstruction for 3D Large Language Models 19 Apr 2024 · 0 repositories · arXiv:2404.13044
-
G-HOP: Generative Hand-Object Prior for Interaction Reconstruction and Grasp Synthesis 18 Apr 2024 · 0 repositories · arXiv:2404.12383
-
Omniview-Tuning: Boosting Viewpoint Invariance of Vision-Language Pre-training Models 18 Apr 2024 · 0 repositories · arXiv:2404.12139
-
The devil is in the object boundary: towards annotation-free instance segmentation using Foundation Models 18 Apr 2024 · 1 repository · arXiv:2404.11957Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
What does CLIP know about peeling a banana? 18 Apr 2024 · 0 repositories · arXiv:2404.12015
-
A Progressive Framework of Vision-language Knowledge Distillation and Alignment for Multilingual Scene 17 Apr 2024 · 0 repositories · arXiv:2404.11249
-
Lightweight Unsupervised Federated Learning with Pretrained Vision Language Model 17 Apr 2024 · 0 repositories · arXiv:2404.11046
-
Optimization of Prompt Learning via Multi-Knowledge Representation for Vision-Language Models 16 Apr 2024 · 0 repositories · arXiv:2404.10357
-
Cross-Modal Self-Training: Aligning Images and Pointclouds to Learn Classification without Labels 15 Apr 2024 · 1 repository · arXiv:2404.10146
-
Evolving Interpretable Visual Classifiers with Large Language Models 15 Apr 2024 · 0 repositories · arXiv:2404.09941
-
Leveraging Temporal Contextualization for Video Action Recognition 15 Apr 2024 · 2 repositories · arXiv:2404.09490Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Photo-Realistic Image Restoration in the Wild with Controlled Vision-Language Models 15 Apr 2024 · 2 repositories · arXiv:2404.09732Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 2 violated, 3 with no contract checked; 3 where Syntology's instrument failed) · 1 unverified (of 9 harvested samples) · 4 pointer-only (licence)
-
RankCLIP: Ranking-Consistent Language-Image Pretraining 15 Apr 2024 · 1 repository · arXiv:2404.09387Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 3 where Syntology's instrument failed) · 3 unverified (of 11 harvested samples) · 11 pointer-only (licence)
-
A Realistic Protocol for Evaluation of Weakly Supervised Object Localization 15 Apr 2024 · 1 repository · arXiv:2404.10034
-
The Devil is in the Few Shots: Iterative Visual Knowledge Completion for Few-shot Learning 15 Apr 2024 · 1 repository · arXiv:2404.09778
-
AMU-Tuning: Effective Logit Bias for CLIP-based Few-shot Learning 13 Apr 2024 · 1 repository · arXiv:2404.08958Syntology official (archive's flag): 4 ran · 4 ran (of which 2 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Understanding Multimodal Deep Neural Networks: A Concept Selection View 13 Apr 2024 · 0 repositories · arXiv:2404.08964
-
Detecting AI-Generated Images via CLIP 12 Apr 2024 · 0 repositories · arXiv:2404.08788
-
Enhancing Traffic Safety with Parallel Dense Video Captioning for End-to-End Event Analysis 12 Apr 2024 · 1 repository · arXiv:2404.08229
-
Generalized Contrastive Learning for Multi-Modal Retrieval and Ranking 12 Apr 2024 · 1 repository · arXiv:2404.08535
-
Improving Continuous Sign Language Recognition with Adapted Image Models 12 Apr 2024 · 1 repository · arXiv:2404.08226
-
Vision-Aware Text Features in Referring Image Segmentation: From Object Understanding to Context Understanding 12 Apr 2024 · 1 repository · arXiv:2404.08590
-
Pay Attention to Your Neighbours: Training-Free Open-Vocabulary Semantic Segmentation 12 Apr 2024 · 1 repository · arXiv:2404.08181Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 3 where Syntology's instrument failed) · 2 unverified (of 6 harvested samples) · 2 pointer-only (licence)
-
Scaling (Down) CLIP: A Comprehensive Analysis of Data, Architecture, and Training Strategies 12 Apr 2024 · 0 repositories · arXiv:2404.08197
-
Semantic Approach to Quantifying the Consistency of Diffusion Model Image Generation 12 Apr 2024 · 1 repository · arXiv:2404.08799
-
Text Prompt with Normality Guidance for Weakly Supervised Video Anomaly Detection 12 Apr 2024 · 0 repositories · arXiv:2404.08531
-
Implicit and Explicit Language Guidance for Diffusion-based Visual Perception 11 Apr 2024 · 0 repositories · arXiv:2404.07600
-
PromptSync: Bridging Domain Gaps in Vision-Language Models through Class-Aware Prototype Alignment and Discrimination 11 Apr 2024 · 0 repositories · arXiv:2404.07520
-
Two Effects, One Trigger: On the Modality Gap, Object Bias, and Information Imbalance in Contrastive Vision-Language Models 11 Apr 2024 · 1 repository · arXiv:2404.07983Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 4 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples)
-
View Selection for 3D Captioning via Diffusion Ranking 11 Apr 2024 · 2 repositories · arXiv:2404.07984Syntology 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
BRAVE: Broadening the visual encoding of vision-language models 10 Apr 2024 · 0 repositories · arXiv:2404.07204
-
O-TALC: Steps Towards Combating Oversegmentation within Online Action Segmentation 10 Apr 2024 · 0 repositories · arXiv:2404.06894
-
PEAVS: Perceptual Evaluation of Audio-Visual Synchrony Grounded in Viewers' Opinion Scores 10 Apr 2024 · 1 repository · arXiv:2404.07336
-
Anchor-based Robust Finetuning of Vision-Language Models 9 Apr 2024 · 0 repositories · arXiv:2404.06244
-
Audio-Visual Generalized Zero-Shot Learning using Pre-Trained Large Multi-Modal Models 9 Apr 2024 · 1 repository · arXiv:2404.06309
-
CLIP-Embed-KD: Computationally Efficient Knowledge Distillation Using Embeddings as Teachers 9 Apr 2024 · 1 repository · arXiv:2404.06170
-
OmniFusion Technical Report 9 Apr 2024 · 0 repositories · arXiv:2404.06212
-
PURE: Turning Polysemantic Neurons Into Pure Features by Identifying Relevant Circuits 9 Apr 2024 · 1 repository · arXiv:2404.06453Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples)
-
Test-Time Adaptation with SaLIP: A Cascade of SAM and CLIP for Zero shot Medical Image Segmentation 9 Apr 2024 · 1 repository · arXiv:2404.06362Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 2 where Syntology's instrument failed) · 3 unverified (of 12 harvested samples) · 5 pointer-only (licence)
-
StyleForge: Enhancing Text-to-Image Synthesis for Any Artistic Styles with Dual Binding 8 Apr 2024 · 0 repositories · arXiv:2404.05256
-
Towards More General Video-based Deepfake Detection through Facial Feature Guided Adaptation for Foundation Model 8 Apr 2024 · 1 repository · arXiv:2404.05583
-
UniMD: Towards Unifying Moment Retrieval and Temporal Action Detection 7 Apr 2024 · 1 repository · arXiv:2404.04933
-
To Cool or not to Cool? Temperature Network Meets Large Foundation Models via DRO 6 Apr 2024 · 1 repository · arXiv:2404.04575Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 3 where Syntology's instrument failed) · 3 unverified (of 8 harvested samples) · 8 pointer-only (licence)
-
Context-Aware Aerial Object Detection: Leveraging Inter-Object and Background Relationships 5 Apr 2024 · 0 repositories · arXiv:2404.04140
-
Diverse and Tailored Image Generation for Zero-shot Multi-label Classification 4 Apr 2024 · 0 repositories · arXiv:2404.03144
-
Is CLIP the main roadblock for fine-grained open-world perception? 4 Apr 2024 · 2 repositories · arXiv:2404.03539Syntology official (archive's flag): 13 ran · 15 ran (of which 0 constructed an object rather than computing a result; 12 with no instrument failure: 0 honoured, 0 violated, 12 with no contract checked; 3 where Syntology's instrument failed) · 2 unverified (of 17 harvested samples) · 17 pointer-only (licence)
-
No "Zero-Shot" Without Exponential Data: Pretraining Concept Frequency Determines Multimodal Model Performance 4 Apr 2024 · 1 repository · arXiv:2404.04125Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 0 where Syntology's instrument failed) · 6 unverified (of 14 harvested samples)
-
OpenNeRF: Open Set 3D Neural Scene Segmentation with Pixel-Wise Features and Rendered Novel Views 4 Apr 2024 · 0 repositories · arXiv:2404.03650
-
Sparse Concept Bottleneck Models: Gumbel Tricks in Contrastive Learning 4 Apr 2024 · 2 repositories · arXiv:2404.03323
-
ASAP: Interpretable Analysis and Summarization of AI-generated Image Patterns at Scale 3 Apr 2024 · 0 repositories · arXiv:2404.02990
-
AWOL: Analysis WithOut synthesis using Language 3 Apr 2024 · 0 repositories · arXiv:2404.03042
-
BCAmirs at SemEval-2024 Task 4: Beyond Words: A Multimodal and Multilingual Exploration of Persuasion in Memes 3 Apr 2024 · 1 repository · arXiv:2404.03022
-
Iterated Learning Improves Compositionality in Large Vision-Language Models 2 Apr 2024 · 0 repositories · arXiv:2404.02145
-
Jailbreaking Prompt Attack: A Controllable Adversarial Attack against Diffusion Models 2 Apr 2024 · 0 repositories · arXiv:2404.02928
-
LP++: A Surprisingly Strong Linear Probe for Few-Shot CLIP 2 Apr 2024 · 1 repository · arXiv:2404.02285Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 2 where Syntology's instrument failed) · 2 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
R^2-Tuning: Efficient Image-to-Video Transfer Learning for Video Temporal Grounding 2 Apr 2024 · 1 repository
-
RAVE: Residual Vector Embedding for CLIP-Guided Backlit Image Enhancement 2 Apr 2024 · 1 repository · arXiv:2404.01889
-
VLRM: Vision-Language Models act as Reward Models for Image Captioning 2 Apr 2024 · 0 repositories · arXiv:2404.01911
-
CLIPtone: Unsupervised Learning for Text-based Image Tone Adjustment 1 Apr 2024 · 0 repositories · arXiv:2404.01123
-
Evaluating Text-to-Visual Generation with Image-to-Text Generation 1 Apr 2024 · 3 repositories · arXiv:2404.01291
-
Perceptogram: Reconstructing Visual Percepts from EEG 1 Apr 2024 · 1 repository · arXiv:2404.01250
-
Meta Episodic learning with Dynamic Task Sampling for CLIP-based Point Cloud Classification 1 Apr 2024 · 0 repositories · arXiv:2404.00857
-
Towards Memorization-Free Diffusion Models 1 Apr 2024 · 0 repositories · arXiv:2404.00922
-
R²-Tuning: Efficient Image-to-Video Transfer Learning for Video Temporal Grounding 31 Mar 2024 · 1 repository · arXiv:2404.00801
-
Training-Free Semantic Segmentation via LLM-Supervision 31 Mar 2024 · 0 repositories · arXiv:2404.00701
-
Unknown Prompt, the only Lacuna: Unveiling CLIP's Potential for Open Domain Generalization 31 Mar 2024 · 1 repository · arXiv:2404.00710
-
CLIP-driven Outliers Synthesis for few-shot OOD detection 30 Mar 2024 · 0 repositories · arXiv:2404.00323
-
Do Vision-Language Models Understand Compound Nouns? 30 Mar 2024 · 1 repository · arXiv:2404.00419Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
FreeSeg-Diff: Training-Free Open-Vocabulary Segmentation with Diffusion Models 29 Mar 2024 · 0 repositories · arXiv:2403.20105
-
Learn "No" to Say "Yes" Better: Improving Vision-Language Models via Negations 29 Mar 2024 · 1 repository · arXiv:2403.20312Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
MedCLIP-SAM: Bridging Text and Image Towards Universal Medical Image Segmentation 29 Mar 2024 · 1 repository · arXiv:2403.20253
-
CLAP4CLIP: Continual Learning with Probabilistic Finetuning for Vision-Language Models 28 Mar 2024 · 1 repository · arXiv:2403.19137Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
Concept-based Analysis of Neural Networks via Vision-Language Models 28 Mar 2024 · 0 repositories · arXiv:2403.19837
-
Model Stock: All we need is just a few fine-tuned models 28 Mar 2024 · 2 repositories · arXiv:2403.19522
-
RH20T-P: A Primitive-Level Robotic Dataset Towards Composable Generalization Agents 28 Mar 2024 · 0 repositories · arXiv:2403.19622
-
Text Data-Centric Image Captioning with Interactive Prompts 28 Mar 2024 · 0 repositories · arXiv:2403.19193
-
Beyond Embeddings: The Promise of Visual Table in Visual Reasoning 27 Mar 2024 · 1 repository · arXiv:2403.18252Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 1 honoured, 2 violated, 3 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 8 harvested samples) · 1 pointer-only (licence)
-
ImageNet-D: Benchmarking Neural Network Robustness on Diffusion Synthetic Object 27 Mar 2024 · 1 repository · arXiv:2403.18775Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 1 honoured, 0 violated, 2 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 6 harvested samples)
-
Language Plays a Pivotal Role in the Object-Attribute Compositional Generalization of CLIP 27 Mar 2024 · 0 repositories · arXiv:2403.18525