Methods › Computer Vision › Vision and Language Pre-Trained Models › CLIP › Papers, page 20
Contrastive Language-Image Pre-training
CLIP
Papers archive 2025-07-28
archive papers tagged: 3,094 · with a code link: 1,617 · where Syntology ran a sample: 649 (554 with a run with no instrument failure, 95 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (649 of 3,094 tagged: 554 with a run with no instrument failure, 95 where every run was a failure of Syntology's instrument)
Page 20 of 31: papers 1,901 to 2,000 of 3,094, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
CP-EB: Talking Face Generation with Controllable Pose and Eye Blinking Embedding 15 Nov 2023 · 0 repositories · arXiv:2311.08673
-
Domain Aligned CLIP for Few-shot Classification 15 Nov 2023 · 0 repositories · arXiv:2311.09191
-
Fast Certification of Vision-Language Models Using Incremental Randomized Smoothing 15 Nov 2023 · 0 repositories · arXiv:2311.09024
-
WildlifeDatasets: An open-source toolkit for animal re-identification 15 Nov 2023 · 2 repositories · arXiv:2311.09118Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 8 harvested samples) · 2 pointer-only (licence)
-
Peer is Your Pillar: A Data-unbalanced Conditional GANs for Few-shot Image Generation 14 Nov 2023 · 0 repositories · arXiv:2311.08217
-
CLiF-VQA: Enhancing Video Quality Assessment by Incorporating High-Level Semantic Information related to Human Feelings 13 Nov 2023 · 0 repositories · arXiv:2311.07090
-
Pretrain like Your Inference: Masked Tuning Improves Zero-Shot Composed Image Retrieval 13 Nov 2023 · 1 repository · arXiv:2311.07622Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Follow-Up Differential Descriptions: Language Models Resolve Ambiguities for Image Classification 10 Nov 2023 · 1 repository · arXiv:2311.07593Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Watermarking Vision-Language Pre-trained Models for Multi-modal Embedding as a Service 10 Nov 2023 · 1 repository · arXiv:2311.05863
-
3DStyle-Diffusion: Pursuing Fine-grained Text-driven 3D Stylization with 2D Diffusion Models 9 Nov 2023 · 1 repository · arXiv:2311.05464
-
GIPCOL: Graph-Injected Soft Prompting for Compositional Zero-Shot Learning 9 Nov 2023 · 1 repository · arXiv:2311.05729Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 5 unverified (of 12 harvested samples) · 12 pointer-only (licence)
-
Language-guided Robot Grasping: CLIP-based Referring Grasp Synthesis in Clutter 9 Nov 2023 · 1 repository · arXiv:2311.05779
-
Enhancing Few-shot CLIP with Semantic-Aware Fine-Tuning 8 Nov 2023 · 0 repositories · arXiv:2311.04464
-
Image-Based Virtual Try-On: A Survey 8 Nov 2023 · 1 repository · arXiv:2311.04811
-
Training CLIP models on Data from Scientific Papers 8 Nov 2023 · 1 repository · arXiv:2311.04711
-
Weakly supervised cross-modal learning in high-content screening 8 Nov 2023 · 1 repository · arXiv:2311.04678Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 11 harvested samples)
-
Can CLIP Help Sound Source Localization? 7 Nov 2023 · 1 repository · arXiv:2311.04066Syntology official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 10 harvested samples) · 10 pointer-only (licence)
-
CLIP Guided Image-perceptive Prompt Learning for Image Enhancement 7 Nov 2023 · 0 repositories · arXiv:2311.03943
-
Meta-Adapter: An Online Few-shot Learner for Vision-Language Model 7 Nov 2023 · 1 repository · arXiv:2311.03774
-
Selective Visual Representations Improve Convergence and Generalization for Embodied AI 7 Nov 2023 · 0 repositories · arXiv:2311.04193
-
Robust Fine-Tuning of Vision-Language Models for Domain Generalization 3 Nov 2023 · 1 repository · arXiv:2311.02236
-
Align Your Prompts: Test-Time Prompting with Distribution Alignment for Zero-Shot Generalization 2 Nov 2023 · 0 repositories · arXiv:2311.01459
-
Learning to Adapt CLIP for Few-Shot Monocular Depth Estimation 2 Nov 2023 · 0 repositories · arXiv:2311.01034
-
Recognize Any Regions 2 Nov 2023 · 1 repository · arXiv:2311.01373Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 3 where Syntology's instrument failed) · 6 unverified (of 11 harvested samples) · 11 pointer-only (licence)
-
CLIP-AD: A Language-Guided Staged Dual-Path Model for Zero-shot Anomaly Detection 1 Nov 2023 · 0 repositories · arXiv:2311.00453
-
Re-Scoring Using Image-Language Similarity for Few-Shot Object Detection 1 Nov 2023 · 1 repository · arXiv:2311.00278
-
ZEETAD: Adapting Pretrained Vision-Language Model for Zero-Shot End-to-End Temporal Action Detection 1 Nov 2023 · 0 repositories · arXiv:2311.00729
-
Class Incremental Learning with Pre-trained Vision-Language Models 31 Oct 2023 · 0 repositories · arXiv:2310.20348
-
Diversity and Diffusion: Observations on Synthetic Image Distributions with Stable Diffusion 31 Oct 2023 · 0 repositories · arXiv:2311.00056
-
Language Guided Visual Question Answering: Elevate Your Multimodal Language Model Using Knowledge-Enriched Prompts 31 Oct 2023 · 1 repository · arXiv:2310.20159Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
MCAD: Multi-teacher Cross-modal Alignment Distillation for efficient image-text retrieval 30 Oct 2023 · 0 repositories · arXiv:2310.19654
-
VideoCrafter1: Open Diffusion Models for High-Quality Video Generation 30 Oct 2023 · 3 repositories · arXiv:2310.19512Syntology community repositories only · 10 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 1 honoured, 0 violated, 8 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 12 harvested samples) · 7 pointer-only (licence)
-
AnomalyCLIP: Object-agnostic Prompt Learning for Zero-shot Anomaly Detection 29 Oct 2023 · 3 repositories · arXiv:2310.18961Syntology official (archive's flag): 12 ran · 19 ran (of which 0 constructed an object rather than computing a result; 14 with no instrument failure: 1 honoured, 1 violated, 12 with no contract checked; 5 where Syntology's instrument failed) · 13 unverified (of 32 harvested samples) · 15 pointer-only (licence)
-
Customize StyleGAN with One Hand Sketch 29 Oct 2023 · 0 repositories · arXiv:2310.18949
-
Text Augmented Spatial-aware Zero-shot Referring Image Segmentation 27 Oct 2023 · 0 repositories · arXiv:2310.18049
-
A Hybrid Graph Network for Complex Activity Detection in Video 26 Oct 2023 · 0 repositories · arXiv:2310.17493
-
Learning Temporal Sentence Grounding From Narrated EgoVideos 26 Oct 2023 · 1 repository · arXiv:2310.17395
-
LP-OVOD: Open-Vocabulary Object Detection by Linear Probing 26 Oct 2023 · 1 repository · arXiv:2310.17109Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Prototypical Contrastive Learning-based CLIP Fine-tuning for Object Re-identification 26 Oct 2023 · 1 repository · arXiv:2310.17218
-
EmoCLIP: A Vision-Language Method for Zero-Shot Video Facial Expression Recognition 25 Oct 2023 · 1 repository · arXiv:2310.16640Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 6 unverified (of 10 harvested samples)
-
Kiki or Bouba? Sound Symbolism in Vision-and-Language Models 25 Oct 2023 · 0 repositories · arXiv:2310.16781
-
Lang3DSG: Language-based contrastive pre-training for 3D Scene Graph prediction 25 Oct 2023 · 0 repositories · arXiv:2310.16494
-
Learning with Noisy Labels Using Collaborative Sample Selection and Contrastive Semi-Supervised Learning 24 Oct 2023 · 0 repositories · arXiv:2310.15533
-
TiC-CLIP: Continual Training of CLIP Models 24 Oct 2023 · 1 repository · arXiv:2310.16226Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 9 harvested samples) · 9 pointer-only (licence)
-
Open-Set Image Tagging with Multi-Grained Text Supervision 23 Oct 2023 · 2 repositories · arXiv:2310.15200Syntology community repositories only · 7 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 8 harvested samples) · 3 pointer-only (licence)
-
Leveraging Image-Text Similarity and Caption Modification for the DataComp Challenge: Filtering Track and BYOD Track 23 Oct 2023 · 0 repositories · arXiv:2310.14581
-
SAM-CLIP: Merging Vision Foundation Models towards Semantic and Spatial Understanding 23 Oct 2023 · 0 repositories · arXiv:2310.15308
-
The BLA Benchmark: Investigating Basic Language Abilities of Pre-Trained Multimodal Models 23 Oct 2023 · 1 repository · arXiv:2310.15061Syntology official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 5 where Syntology's instrument failed) · 1 unverified (of 12 harvested samples)
-
Unleashing the potential of prompt engineering for large language models 23 Oct 2023 · 0 repositories · arXiv:2310.14735
-
Unveiling the Power of CLIP in Unsupervised Visible-Infrared Person Re-Identification 23 Oct 2023 · 1 repository
-
One-for-All: Towards Universal Domain Translation with a Single StyleGAN 22 Oct 2023 · 0 repositories · arXiv:2310.14222
-
CLIP meets Model Zoo Experts: Pseudo-Supervision for Visual Enhancement 21 Oct 2023 · 0 repositories · arXiv:2310.14108
-
CAPIVARA: Cost-Efficient Approach for Improving Multilingual CLIP Performance on Low-Resource Languages 20 Oct 2023 · 1 repository · arXiv:2310.13683
-
Localizing and Editing Knowledge in Text-to-Image Generative Models 20 Oct 2023 · 0 repositories · arXiv:2310.13730
-
On the Language Encoder of Contrastive Cross-modal Models 20 Oct 2023 · 0 repositories · arXiv:2310.13267
-
Reference-based Restoration of Digitized Analog Videotapes 20 Oct 2023 · 2 repositories · arXiv:2310.14926
-
SILC: Improving Vision Language Pretraining with Self-Distillation 20 Oct 2023 · 0 repositories · arXiv:2310.13355
-
Interpreting CLIP: Insights on the Robustness to ImageNet Distribution Shifts 19 Oct 2023 · 0 repositories · arXiv:2310.13040
-
Vision-Language Models are Zero-Shot Reward Models for Reinforcement Learning 19 Oct 2023 · 1 repository · arXiv:2310.12921Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 5 harvested samples)
-
Evaluating the Fairness of Discriminative Foundation Models in Computer Vision 18 Oct 2023 · 1 repository · arXiv:2310.11867
-
On the use of Vision-Language models for Visual Sentiment Analysis: a study on CLIP 18 Oct 2023 · 2 repositories · arXiv:2310.12062
-
Combating Label Noise With A General Surrogate Model For Sample Selection 16 Oct 2023 · 0 repositories · arXiv:2310.10463
-
Interpreting and Controlling Vision Foundation Models via Text Explanations 16 Oct 2023 · 1 repository · arXiv:2310.10591
-
Prompting Scientific Names for Zero-Shot Species Recognition 15 Oct 2023 · 0 repositories · arXiv:2310.09929
-
Does CLIP's Generalization Performance Mainly Stem from High Train-Test Similarity? 14 Oct 2023 · 1 repository · arXiv:2310.09562Syntology official: harvested, nothing ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 3 harvested samples) · 1 pointer-only (licence)
-
Efficient Model-Agnostic Multi-Group Equivariant Networks 14 Oct 2023 · 0 repositories · arXiv:2310.09675
-
Extending Multi-modal Contrastive Representations 13 Oct 2023 · 1 repository · arXiv:2310.08884Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models 13 Oct 2023 · 1 repository · arXiv:2310.08825
-
Incremental Object Detection with CLIP 13 Oct 2023 · 0 repositories · arXiv:2310.08815
-
EasyGen: Easing Multimodal Generation with BiDiffuser and LLMs 13 Oct 2023 · 1 repository · arXiv:2310.08949
-
Vision-by-Language for Training-Free Compositional Image Retrieval 13 Oct 2023 · 1 repository · arXiv:2310.09291Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 1 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 8 harvested samples) · 2 pointer-only (licence)
-
Defending Our Privacy With Backdoors 12 Oct 2023 · 1 repository · arXiv:2310.08320
-
DeltaSpace: A Semantic-aligned Feature Space for Flexible Text-guided Image Editing 12 Oct 2023 · 1 repository · arXiv:2310.08785
-
Leveraging Vision-Language Models for Improving Domain Generalization in Image Classification 12 Oct 2023 · 1 repository · arXiv:2310.08255Syntology official (archive's flag): 12 ran · 12 ran (of which 0 constructed an object rather than computing a result; 11 with no instrument failure: 1 honoured, 0 violated, 10 with no contract checked; 1 where Syntology's instrument failed) · 3 unverified (of 15 harvested samples) · 2 pointer-only (licence)
-
Generalized Logit Adjustment: Calibrating Fine-tuned Models by Removing Label Bias in Foundation Models 12 Oct 2023 · 2 repositories · arXiv:2310.08106
-
Mapping Memes to Words for Multimodal Hateful Meme Classification 12 Oct 2023 · 1 repository · arXiv:2310.08368Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples)
-
Visual Data-Type Understanding does not emerge from Scaling Vision-Language Models 12 Oct 2023 · 1 repository · arXiv:2310.08577
-
CLIP for Lightweight Semantic Segmentation 11 Oct 2023 · 0 repositories · arXiv:2310.07394
-
ConditionVideo: Training-Free Condition-Guided Text-to-Video Generation 11 Oct 2023 · 1 repository · arXiv:2310.07697
-
VeCLIP: Improving CLIP Training via Visual-enriched Captions 11 Oct 2023 · 1 repository · arXiv:2310.07699
-
AutoAD II: The Sequel -- Who, When, and What in Movie Audio Description 10 Oct 2023 · 0 repositories · arXiv:2310.06838
-
Blind Dates: Examining the Expression of Temporality in Historical Photographs 10 Oct 2023 · 0 repositories · arXiv:2310.06633
-
Cross-modal Cognitive Consensus guided Audio-Visual Segmentation 10 Oct 2023 · 1 repository · arXiv:2310.06259Syntology official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 12 harvested samples) · 12 pointer-only (licence)
-
Robustness May be More Brittle than We Think under Different Degrees of Distribution Shifts 10 Oct 2023 · 0 repositories · arXiv:2310.06622
-
A General Protocol to Probe Large Vision Models for 3D Physical Understanding 10 Oct 2023 · 1 repository · arXiv:2310.06836
-
Interpreting CLIP's Image Representation via Text-Based Decomposition 9 Oct 2023 · 1 repository · arXiv:2310.05916Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Text-driven Prompt Generation for Vision-Language Models in Federated Learning 9 Oct 2023 · 0 repositories · arXiv:2310.06123
-
Building an Open-Vocabulary Video CLIP Model with Better Architectures, Optimization and Data 8 Oct 2023 · 1 repository · arXiv:2310.05010
-
Compositional Semantics for Open Vocabulary Spatio-semantic Representations 8 Oct 2023 · 0 repositories · arXiv:2310.04981
-
GMMFormer: Gaussian-Mixture-Model Based Transformer for Efficient Partially Relevant Video Retrieval 8 Oct 2023 · 1 repository · arXiv:2310.05195Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 1 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
Symmetrical Linguistic Feature Distillation with CLIP for Scene Text Recognition 8 Oct 2023 · 1 repository · arXiv:2310.04999
-
Understanding the Robustness of Multi-modal Contrastive Learning to Distribution Shift 8 Oct 2023 · 0 repositories · arXiv:2310.04971
-
DeVAn: Dense Video Annotation for Video-Language Models 8 Oct 2023 · 1 repository · arXiv:2310.05060
-
Better Safe than Sorry: Pre-training CLIP against Targeted Data Poisoning and Backdoor Attacks 5 Oct 2023 · 1 repository · arXiv:2310.05862Syntology official (archive's flag): 4 ran · 5 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 1 where Syntology's instrument failed) · 5 unverified (of 10 harvested samples) · 10 pointer-only (licence)
-
Investigating the Limitation of CLIP Models: The Worst-Performing Categories 5 Oct 2023 · 0 repositories · arXiv:2310.03324
-
Kandinsky: an Improved Text-to-Image Synthesis with Image Prior and Latent Diffusion 5 Oct 2023 · 1 repository · arXiv:2310.03502Syntology official (archive's flag): 3 ran · 12 ran (of which 1 constructed an object rather than computing a result; 7 with no instrument failure: 3 honoured, 0 violated, 4 with no contract checked; 5 where Syntology's instrument failed) · 6 unverified (of 18 harvested samples) · 10 pointer-only (licence)
-
Delving into CLIP latent space for Video Anomaly Recognition 4 Oct 2023 · 1 repository · arXiv:2310.02835Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Learning to Prompt Your Domain for Vision-Language Models 4 Oct 2023 · 0 repositories · arXiv:2310.03103
-
Kosmos-G: Generating Images in Context with Multimodal Large Language Models 4 Oct 2023 · 1 repository · arXiv:2310.02992
-
Mending of Spatio-Temporal Dependencies in Block Adjacency Matrix 4 Oct 2023 · 0 repositories · arXiv:2310.02606