Methods › Computer Vision › Vision and Language Pre-Trained Models › CLIP › Papers, page 29
Contrastive Language-Image Pre-training
CLIP
Papers archive 2025-07-28
archive papers tagged: 3,094 · with a code link: 1,617 · where Syntology ran a sample: 649 (554 with a run with no instrument failure, 95 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (649 of 3,094 tagged: 554 with a run with no instrument failure, 95 where every run was a failure of Syntology's instrument)
Page 29 of 31: papers 2,801 to 2,900 of 3,094, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Diffusion-based Image Translation using Disentangled Style and Content Representation 30 Sep 2022 · 1 repository · arXiv:2209.15264Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample)
-
Linearly Mapping from Image to Text Space 30 Sep 2022 · 2 repositories · arXiv:2209.15162Syntology official (archive's flag): 13 ran · 13 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 3 where Syntology's instrument failed) · 1 unverified (of 14 harvested samples)
-
SmallCap: Lightweight Image Captioning Prompted with Retrieval Augmentation 30 Sep 2022 · 1 repository · arXiv:2209.15323
-
Understanding Pure CLIP Guidance for Voxel Grid NeRF Models 30 Sep 2022 · 0 repositories · arXiv:2209.15172
-
REST: REtrieve & Self-Train for generative action recognition 29 Sep 2022 · 0 repositories · arXiv:2209.15000
-
360FusionNeRF: Panoramic Neural Radiance Fields with Joint Guidance 28 Sep 2022 · 1 repository · arXiv:2209.14265
-
CALIP: Zero-Shot Enhancement of CLIP with Parameter-free Attention 28 Sep 2022 · 1 repository · arXiv:2209.14169
-
Dynamic MDETR: A Dynamic Multimodal Transformer Decoder for Visual Grounding 28 Sep 2022 · 0 repositories · arXiv:2209.13959
-
NEURAL MARIONETTE: A Transformer-based Multi-action Human Motion Synthesis System 27 Sep 2022 · 0 repositories · arXiv:2209.13204
-
Collaboration of Pre-trained Models Makes Better Few-shot Learner 25 Sep 2022 · 0 repositories · arXiv:2209.12255
-
NamedMask: Distilling Segmenters from Complementary Foundation Models 22 Sep 2022 · 1 repository · arXiv:2209.11228Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples)
-
Continual VQA for Disaster Response Systems 21 Sep 2022 · 1 repository · arXiv:2209.10320
-
GAMA: Generative Adversarial Multi-Object Scene Attacks 20 Sep 2022 · 1 repository · arXiv:2209.09502Syntology official (archive's flag): 2 ran · 2 ran (of which 2 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified; every one of the 2 samples that ran constructed an object rather than computing a result (of 2 harvested samples)
-
Text2Light: Zero-Shot Text-Driven HDR Panorama Generation 20 Sep 2022 · 1 repository · arXiv:2209.09898
-
Effective Adaptation in Multi-Task Co-Training for Unified Autonomous Driving 19 Sep 2022 · 0 repositories · arXiv:2209.08953
-
Exploiting Cultural Biases via Homoglyphs in Text-to-Image Synthesis 19 Sep 2022 · 2 repositories · arXiv:2209.08891Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Does CLIP Know My Face? 15 Sep 2022 · 3 repositories · arXiv:2209.07341Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Exploring Visual Interpretability for Contrastive Language-Image Pre-training 15 Sep 2022 · 1 repository · arXiv:2209.07046
-
Test-Time Prompt Tuning for Zero-Shot Generalization in Vision-Language Models 15 Sep 2022 · 1 repository · arXiv:2209.07511Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 2 unverified (of 6 harvested samples)
-
CLIP-ViP: Adapting Pre-trained Image-Text Model to Video-Language Representation Alignment 14 Sep 2022 · 1 repository · arXiv:2209.06430Syntology official (archive's flag): 4 ran · 4 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Generative Visual Prompt: Unifying Distributional Control of Pre-Trained Generative Models 14 Sep 2022 · 1 repository · arXiv:2209.06970
-
VL-Taboo: An Analysis of Attribute-based Zero-shot Capabilities of Vision-Language Models 12 Sep 2022 · 1 repository · arXiv:2209.06103
-
ISS: Image as Stepping Stone for Text-Guided 3D Shape Generation 9 Sep 2022 · 2 repositories · arXiv:2209.04145
-
Text-Free Learning of a Natural Language Interface for Pretrained Face Generators 8 Sep 2022 · 1 repository · arXiv:2209.03953
-
AI Illustrator: Translating Raw Descriptions into Images by Prompt-based Cross-Modal Generation 7 Sep 2022 · 1 repository · arXiv:2209.03160
-
Every picture tells a story: Image-grounded controllable stylistic story generation 4 Sep 2022 · 0 repositories · arXiv:2209.01638
-
Video-Guided Curriculum Learning for Spoken Video Grounding 1 Sep 2022 · 1 repository · arXiv:2209.00277Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 1 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 1 pointer-only (licence)
-
Zero-Shot Multi-Modal Artist-Controlled Retrieval and Exploration of 3D Object Sets 1 Sep 2022 · 0 repositories · arXiv:2209.00682
-
DetailCLIP: Injecting Image Details into CLIP's Feature Space 31 Aug 2022 · 0 repositories · arXiv:2208.14649
-
TCAM: Temporal Class Activation Maps for Object Localization in Weakly-Labeled Unconstrained Videos 30 Aug 2022 · 1 repository · arXiv:2208.14542
-
Efficient Vision-Language Pretraining with Visual Concepts and Hierarchical Alignment 29 Aug 2022 · 1 repository · arXiv:2208.13628Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
LogicRank: Logic Induced Reranking for Generative Text-to-Image Systems 29 Aug 2022 · 0 repositories · arXiv:2208.13518
-
Cross-Lingual Cross-Modal Retrieval with Noise-Robust Learning 26 Aug 2022 · 1 repository · arXiv:2208.12526
-
PromptFL: Let Federated Participants Cooperatively Learn Prompts Instead of Models -- Federated Learning in Age of Foundation Model 24 Aug 2022 · 0 repositories · arXiv:2208.11625
-
Semantic-Enhanced Image Clustering 21 Aug 2022 · 0 repositories · arXiv:2208.09849
-
Dance Style Transfer with Cross-modal Transformer 19 Aug 2022 · 0 repositories · arXiv:2208.09406
-
Open-Vocabulary Universal Image Segmentation with MaskCLIP 18 Aug 2022 · 1 repository · arXiv:2208.08984Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Dual Modality Prompt Tuning for Vision-Language Pre-Trained Model 17 Aug 2022 · 1 repository · arXiv:2208.08340
-
Leveraging Endo- and Exo-Temporal Regularization for Black-box Video Domain Adaptation 10 Aug 2022 · 0 repositories · arXiv:2208.05187
-
Patching open-vocabulary models by interpolating weights 10 Aug 2022 · 1 repository · arXiv:2208.05592Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 3 where Syntology's instrument failed) · 3 unverified (of 8 harvested samples) · 5 pointer-only (licence)
-
Quality Not Quantity: On the Interaction between Dataset Design and Robustness of CLIP 10 Aug 2022 · 1 repository · arXiv:2208.05516
-
Distinctive Image Captioning via CLIP Guided Group Optimization 8 Aug 2022 · 0 repositories · arXiv:2208.04254
-
Weakly Supervised Online Action Detection for Infant General Movements 7 Aug 2022 · 1 repository · arXiv:2208.03648
-
Frozen CLIP Models are Efficient Video Learners 6 Aug 2022 · 2 repositories · arXiv:2208.03550Syntology official (archive's flag): 5 ran · 5 ran (of which 3 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 5 unverified (of 10 harvested samples) · 10 pointer-only (licence)
-
A Novel Enhanced Convolution Neural Network with Extreme Learning Machine: Facial Emotional Recognition in Psychology Practices 5 Aug 2022 · 0 repositories · arXiv:2208.02953
-
A Sketch Is Worth a Thousand Words: Image Retrieval with Text and Sketch 5 Aug 2022 · 0 repositories · arXiv:2208.03354
-
MAFW: A Large-scale, Multi-modal, Compound Affective Database for Dynamic Facial Expression Recognition in the Wild 1 Aug 2022 · 0 repositories · arXiv:2208.00847
-
Curriculum Learning for Data-Efficient Vision-Language Alignment 29 Jul 2022 · 0 repositories · arXiv:2207.14525
-
Learning Visual Representation from Modality-Shared Contrastive Language-Image Pre-training 26 Jul 2022 · 1 repository · arXiv:2207.12661
-
Text-Guided Synthesis of Artistic Images with Retrieval-Augmented Diffusion Models 26 Jul 2022 · 1 repository · arXiv:2207.13038
-
Contrastive Learning for Interactive Recommendation in Fashion 25 Jul 2022 · 0 repositories · arXiv:2207.12033
-
Exploring CLIP for Assessing the Look and Feel of Images 25 Jul 2022 · 1 repository · arXiv:2207.12396
-
Learning Dynamic Facial Radiance Fields for Few-Shot Talking Head Synthesis 24 Jul 2022 · 1 repository · arXiv:2207.11770Syntology official (archive's flag): 1 ran · 1 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified; the one sample that ran constructed an object rather than computing a result (of 1 harvested sample)
-
Weakly-Supervised Temporal Action Detection for Fine-Grained Videos with Hierarchical Atomic Actions 24 Jul 2022 · 1 repository · arXiv:2207.11805
-
Generative Artisan: A Semantic-Aware and Controllable CLIPstyler 23 Jul 2022 · 0 repositories · arXiv:2207.11598
-
Robots Enact Malignant Stereotypes 23 Jul 2022 · 0 repositories · arXiv:2207.11569
-
Semantic Abstraction: Open-World 3D Scene Understanding from 2D Vision-Language Models 23 Jul 2022 · 1 repository · arXiv:2207.11514Syntology official (archive's flag): 15 ran · 15 ran (of which 0 constructed an object rather than computing a result; 15 with no instrument failure: 1 honoured, 0 violated, 14 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 18 harvested samples) · 4 pointer-only (licence)
-
DeVIS: Making Deformable Transformers Work for Video Instance Segmentation 22 Jul 2022 · 1 repository · arXiv:2207.11103Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 1 honoured, 0 violated, 7 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 11 harvested samples) · 11 pointer-only (licence)
-
Zero-Shot Video Captioning with Evolving Pseudo-Tokens 22 Jul 2022 · 1 repository · arXiv:2207.11100
-
Don't Stop Learning: Towards Continual Learning for the CLIP Model 19 Jul 2022 · 0 repositories · arXiv:2207.09248
-
Tip-Adapter: Training-free Adaption of CLIP for Few-shot Classification 19 Jul 2022 · 3 repositories · arXiv:2207.09519Syntology official (archive's flag): 1 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 2 pointer-only (licence)
-
Towards Diverse and Faithful One-shot Adaption of Generative Adversarial Networks 18 Jul 2022 · 1 repository · arXiv:2207.08736Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 1 violated, 2 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples) · 1 pointer-only (licence)
-
Zero-Shot Temporal Action Detection via Vision-Language Prompting 17 Jul 2022 · 1 repository · arXiv:2207.08184
-
Is a Caption Worth a Thousand Images? A Controlled Study for Representation Learning 15 Jul 2022 · 0 repositories · arXiv:2207.07635
-
Contrastive Adapters for Foundation Model Group Robustness 14 Jul 2022 · 0 repositories · arXiv:2207.07180
-
Progressively-connected Light Field Network for Efficient View Synthesis 10 Jul 2022 · 0 repositories · arXiv:2207.04465
-
Models Out of Line: A Fourier Lens on Distribution Shift Robustness 8 Jul 2022 · 0 repositories · arXiv:2207.04075
-
Bridging the Gap between Object and Image-level Representations for Open-Vocabulary Detection 7 Jul 2022 · 1 repository · arXiv:2207.03482Syntology 3 ran (of which 1 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples)
-
GAMa: Cross-view Video Geo-localization 6 Jul 2022 · 1 repository · arXiv:2207.02431
-
Towards Counterfactual Image Manipulation via CLIP 6 Jul 2022 · 1 repository · arXiv:2207.02812Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
LaTeRF: Label and Text Driven Object Radiance Fields 4 Jul 2022 · 0 repositories · arXiv:2207.01583
-
Can Language Understand Depth? 3 Jul 2022 · 1 repository · arXiv:2207.01077Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
American == White in Multimodal Language-and-Image AI 1 Jul 2022 · 0 repositories · arXiv:2207.00691
-
ReLER@ZJU-Alibaba Submission to the Ego4D Natural Language Queries Challenge 2022 1 Jul 2022 · 1 repository · arXiv:2207.00383Syntology official (archive's flag): 17 ran · 17 ran (of which 0 constructed an object rather than computing a result; 16 with no instrument failure: 3 honoured, 0 violated, 13 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 19 harvested samples) · 4 pointer-only (licence)
-
Video + CLIP Baseline for Ego4D Long-term Action Anticipation 1 Jul 2022 · 1 repository · arXiv:2207.00579Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 9 harvested samples) · 1 pointer-only (licence)
-
VL-CheckList: Evaluating Pre-trained Vision-Language Models with Objects, Attributes and Relations 1 Jul 2022 · 1 repository · arXiv:2207.00221
-
Normalized/Clipped SGD with Perturbation for Differentially Private Non-Convex Optimization 27 Jun 2022 · 0 repositories · arXiv:2206.13033
-
Text-Driven Stylization of Video Objects 24 Jun 2022 · 0 repositories · arXiv:2206.12396
-
ProtoCLIP: Prototypical Contrastive Language Image Pretraining 22 Jun 2022 · 1 repository · arXiv:2206.10996
-
Beyond Uniform Lipschitz Condition in Differentially Private Optimization 21 Jun 2022 · 0 repositories · arXiv:2206.10713
-
Conditioned and Composed Image Retrieval Combining and Partially Fine-Tuning CLIP-Based Features 19 Jun 2022 · 2 repositories
-
What is Where by Looking: Weakly-Supervised Open-World Phrase-Grounding without Text Inputs 19 Jun 2022 · 1 repository · arXiv:2206.09358
-
Entity-Graph Enhanced Cross-Modal Pretraining for Instance-level Product Retrieval 17 Jun 2022 · 0 repositories · arXiv:2206.08842
-
Multi-Contextual Predictions with Vision Transformer for Video Anomaly Detection 17 Jun 2022 · 0 repositories · arXiv:2206.08568
-
Know your audience: specializing grounded language models with listener subtraction 16 Jun 2022 · 0 repositories · arXiv:2206.08349
-
MixGen: A New Multi-Modal Data Augmentation 16 Jun 2022 · 1 repository · arXiv:2206.08358Syntology official: harvested, nothing ran · 0 ran · 2 unverified (of 2 harvested samples)
-
Disentangling visual and written concepts in CLIP 15 Jun 2022 · 0 repositories · arXiv:2206.07835
-
ReCo: Retrieve and Co-segment for Zero-shot Transfer 14 Jun 2022 · 2 repositories · arXiv:2206.07045
-
Transductive CLIP with Class-Conditional Contrastive Learning 13 Jun 2022 · 0 repositories · arXiv:2206.06177
-
Less Is More: Linear Layers on CLIP Features as Powerful VizWiz Model 10 Jun 2022 · 0 repositories · arXiv:2206.05281
-
CLIP-Actor: Text-Driven Recommendation and Stylization for Animating Human Meshes 9 Jun 2022 · 1 repository · arXiv:2206.04382Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample)
-
Masked Unsupervised Self-training for Label-free Image Classification 7 Jun 2022 · 1 repository · arXiv:2206.02967Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 1 honoured, 0 violated, 3 with no contract checked; 3 where Syntology's instrument failed) · 4 unverified (of 11 harvested samples) · 4 pointer-only (licence)
-
OrdinalCLIP: Learning Rank Prompts for Language-Guided Ordinal Regression 6 Jun 2022 · 1 repository · arXiv:2206.02338
-
ContraCLIP: Interpretable GAN generation driven by pairs of contrasting sentences 5 Jun 2022 · 1 repository · arXiv:2206.02104
-
Recurrent Video Restoration Transformer with Guided Deformable Attention 5 Jun 2022 · 4 repositories · arXiv:2206.02146Syntology official (archive's flag): 4 ran · 12 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 7 where Syntology's instrument failed) · 2 unverified (of 14 harvested samples) · 7 pointer-only (licence)
-
Delving into the Openness of CLIP 4 Jun 2022 · 1 repository · arXiv:2206.01986Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 3 where Syntology's instrument failed) · 5 unverified (of 10 harvested samples) · 3 pointer-only (licence)
-
Prefix Conditioning Unifies Language and Label Supervision 2 Jun 2022 · 0 repositories · arXiv:2206.01125
-
CLIP4IDC: CLIP for Image Difference Captioning 1 Jun 2022 · 1 repository · arXiv:2206.00629
-
Prompt-aligned Gradient for Prompt Tuning 30 May 2022 · 1 repository · arXiv:2205.14865
-
CyCLIP: Cyclic Contrastive Language-Image Pretraining 28 May 2022 · 1 repository · arXiv:2205.14459Syntology official (archive's flag): 5 ran · 5 ran (of which 4 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 7 harvested samples) · 7 pointer-only (licence)