Methods › Computer Vision › Vision and Language Pre-Trained Models › CLIP › Papers, page 31
Contrastive Language-Image Pre-training
CLIP
Papers archive 2025-07-28
archive papers tagged: 3,094 · with a code link: 1,617 · where Syntology ran a sample: 649 (554 with a run with no instrument failure, 95 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (649 of 3,094 tagged: 554 with a run with no instrument failure, 95 where every run was a failure of Syntology's instrument)
Page 31 of 31: papers 3,001 to 3,094 of 3,094, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
ImaginE: An Imagination-Based Automatic Evaluation Metric for Natural Language Generation 17 Dec 2021 · 0 repositories
-
Soundify: Matching Sound Effects to Video 17 Dec 2021 · 0 repositories · arXiv:2112.09726
-
Contrastive Vision-Language Pre-training with Limited Resources 17 Dec 2021 · 1 repository · arXiv:2112.09331
-
RegionCLIP: Region-based Language-Image Pretraining 16 Dec 2021 · 1 repository · arXiv:2112.09106
-
Twitter-COMMs: Detecting Climate, COVID, and Military Multimodal Misinformation 16 Dec 2021 · 1 repository · arXiv:2112.08594
-
CLIP-Lite: Information Efficient Visual Representation Learning with Language Supervision 14 Dec 2021 · 1 repository · arXiv:2112.07133Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 8 harvested samples) · 1 pointer-only (licence)
-
Multimodal neural networks better explain multivoxel patterns in the hippocampus 11 Dec 2021 · 1 repository · arXiv:2201.11517
-
CLIP2StyleGAN: Unsupervised Extraction of StyleGAN Edit Directions 9 Dec 2021 · 0 repositories · arXiv:2112.05219
-
CMA-CLIP: Cross-Modality Attention CLIP for Image-Text Classification 7 Dec 2021 · 0 repositories · arXiv:2112.03562
-
Embedding Arithmetic of Multimodal Queries for Image Retrieval 6 Dec 2021 · 0 repositories · arXiv:2112.03162
-
Text2Mesh: Text-Driven Neural Stylization for Meshes 6 Dec 2021 · 1 repository · arXiv:2112.03221Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample)
-
PointCLIP: Point Cloud Understanding by CLIP 4 Dec 2021 · 2 repositories · arXiv:2112.02413
-
VT-CLIP: Enhancing Vision-Language Models with Visual-guided Texts 4 Dec 2021 · 0 repositories · arXiv:2112.02399
-
Extract Free Dense Labels from CLIP 2 Dec 2021 · 1 repository · arXiv:2112.01071
-
DenseCLIP: Language-Guided Dense Prediction with Context-Aware Prompting 2 Dec 2021 · 1 repository · arXiv:2112.01518
-
FuseDream: Training-Free Text-to-Image Generation with Improved CLIP+GAN Space Optimization 2 Dec 2021 · 1 repository · arXiv:2112.01573
-
Zero-Shot Text-Guided Object Generation with Dream Fields 2 Dec 2021 · 4 repositories · arXiv:2112.01455Syntology community repositories only · 9 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 10 harvested samples)
-
CLIPstyler: Image Style Transfer with a Single Text Condition 1 Dec 2021 · 3 repositories · arXiv:2112.00374Syntology official (archive's flag): 12 ran · 13 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 5 where Syntology's instrument failed) · 0 unverified (of 13 harvested samples)
-
MAD: A Scalable Dataset for Language Grounding in Videos from Movie Audio Descriptions 1 Dec 2021 · 1 repository · arXiv:2112.00431Syntology official: no sample here; runs from other or unrecorded repositories · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
An implementation of the "Guess who?" game using CLIP 30 Nov 2021 · 1 repository · arXiv:2112.00599
-
CLIP Meets Video Captioning: Concept-Aware Representation Learning Does Matter 30 Nov 2021 · 1 repository · arXiv:2111.15162
-
Blended Diffusion for Text-driven Editing of Natural Images 29 Nov 2021 · 1 repository · arXiv:2111.14818Syntology official (archive's flag): 15 ran · 16 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 4 honoured, 0 violated, 6 with no contract checked; 6 where Syntology's instrument failed) · 6 unverified (of 22 harvested samples) · 13 pointer-only (licence)
-
LAFITE: Towards Language-Free Training for Text-to-Image Generation 27 Nov 2021 · 3 repositories · arXiv:2111.13792Syntology official (archive's flag): 3 ran · 12 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 2 honoured, 0 violated, 8 with no contract checked; 2 where Syntology's instrument failed) · 6 unverified (of 18 harvested samples) · 3 pointer-only (licence)
-
Predict, Prevent, and Evaluate: Disentangled Text-Driven Image Manipulation Empowered by Pre-Trained Vision-Language Model 26 Nov 2021 · 1 repository · arXiv:2111.13333
-
Domain Prompt Learning for Efficiently Adapting CLIP to Unseen Domains 25 Nov 2021 · 1 repository · arXiv:2111.12853Syntology official (archive's flag): 12 ran · 12 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 1 violated, 7 with no contract checked; 4 where Syntology's instrument failed) · 5 unverified (of 17 harvested samples) · 17 pointer-only (licence)
-
Florence: A New Foundation Model for Computer Vision 22 Nov 2021 · 2 repositories · arXiv:2111.11432
-
Combined Scaling for Zero-shot Transfer Learning 19 Nov 2021 · 0 repositories · arXiv:2111.10050
-
ClipCap: CLIP Prefix for Image Captioning 18 Nov 2021 · 4 repositories · arXiv:2111.09734Syntology official (archive's flag): 5 ran · 6 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 6 harvested samples) · 2 pointer-only (licence)
-
Simple but Effective: CLIP Embeddings for Embodied AI 18 Nov 2021 · 2 repositories · arXiv:2111.09888Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples)
-
Image Retrieval from Contextual Descriptions 16 Nov 2021 · 0 repositories
-
Multi-Granularity Contrastive Knowledge Distillation for Multimodal Named Entity Recognition 16 Nov 2021 · 0 repositories
-
Probing the Prompting of CLIP on Human Faces 16 Nov 2021 · 0 repositories
-
ReCLIP: A Strong Zero-Shot Baseline for Referring Expression Comprehension 16 Nov 2021 · 0 repositories
-
Visual-Language Navigation Pretraining via Prompt-based Environmental Self-exploration 16 Nov 2021 · 0 repositories
-
Zero-Shot Visual Grounding of Referring Utterances in Dialogue 16 Nov 2021 · 0 repositories
-
CoLLIE: Continual Learning of Language Grounding from Language-Image Embeddings 15 Nov 2021 · 1 repository · arXiv:2111.07993
-
Scaling Law for Recommendation Models: Towards General-purpose User Representations 15 Nov 2021 · 0 repositories · arXiv:2111.11294
-
Evolving Evocative 2D Views of Generated 3D Objects 8 Nov 2021 · 0 repositories · arXiv:2111.04839
-
Tip-Adapter: Training-free CLIP-Adapter for Better Vision-Language Modeling 6 Nov 2021 · 1 repository · arXiv:2111.03930
-
StyleCLIPDraw: Coupling Content and Style in Text-to-Drawing Synthesis 4 Nov 2021 · 1 repository · arXiv:2111.03133Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs 3 Nov 2021 · 3 repositories · arXiv:2111.02114Syntology official: harvested for another paper · 0 ran · 1 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Image-Based CLIP-Guided Essence Transfer 24 Oct 2021 · 1 repository · arXiv:2110.12427Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
CLOOB: Modern Hopfield Networks with InfoLOOB Outperform CLIP 21 Oct 2021 · 1 repository · arXiv:2110.11316Syntology official (archive's flag): 4 ran · 7 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 6 where Syntology's instrument failed) · 4 unverified (of 11 harvested samples) · 4 pointer-only (licence)
-
Wav2CLIP: Learning Robust Audio Representations From CLIP 21 Oct 2021 · 1 repository · arXiv:2110.11499Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
CIM-PPO:Proximal Policy Optimization with Liu-Correntropy Induced Metric 20 Oct 2021 · 0 repositories · arXiv:2110.10522
-
Risks of AI Foundation Models in Education 19 Oct 2021 · 0 repositories · arXiv:2110.10024
-
Seeing things or seeing scenes: Investigating the capabilities of V&L models to align scene descriptions to images 16 Oct 2021 · 0 repositories
-
Mind the Gap: Domain Gap Control for Single Shot Domain Adaptation for Generative Adversarial Networks 15 Oct 2021 · 2 repositories · arXiv:2110.08398
-
Inverse Problems Leveraging Pre-trained Contrastive Representations 14 Oct 2021 · 1 repository · arXiv:2110.07439
-
Scaling Laws for the Few-Shot Adaptation of Pre-trained Image Classifiers 13 Oct 2021 · 0 repositories · arXiv:2110.06990
-
CLIP4Caption ++: Multi-CLIP for Video Caption 11 Oct 2021 · 0 repositories · arXiv:2110.05204
-
Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training Paradigm 11 Oct 2021 · 4 repositories · arXiv:2110.05208Syntology community repositories only · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 1 honoured, 2 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Inferring Offensiveness In Images From Natural Language Supervision 8 Oct 2021 · 1 repository · arXiv:2110.04222
-
CLIP-Forge: Towards Zero-Shot Text-to-Shape Generation 6 Oct 2021 · 1 repository · arXiv:2110.02624
-
Continual Learning Using Pseudo-Replay via Latent Space Sampling 29 Sep 2021 · 0 repositories
-
Evaluating Language-biased image classification based on semantic compositionality 29 Sep 2021 · 0 repositories
-
How Much Can CLIP Benefit Vision-and-Language Tasks? 29 Sep 2021 · 0 repositories
-
Learning Visual-Linguistic Adequacy, Fidelity, and Fluency for Novel Object Captioning 29 Sep 2021 · 0 repositories
-
MA-CLIP: Towards Modality-Agnostic Contrastive Language-Image Pre-training 29 Sep 2021 · 0 repositories
-
Zero-Shot Reward Specification via Grounded Natural Language 29 Sep 2021 · 0 repositories
-
ClipMatrix: Text-controlled Creation of 3D Textured Meshes 27 Sep 2021 · 1 repository · arXiv:2109.12922
-
CLIPort: What and Where Pathways for Robotic Manipulation 24 Sep 2021 · 1 repository · arXiv:2109.12098
-
Zero-shot Object Detection Through Vision-Language Embedding Alignment 24 Sep 2021 · 1 repository · arXiv:2109.12066
-
What Vision-Language Models `See' when they See Scenes 15 Sep 2021 · 0 repositories · arXiv:2109.07301
-
EfficientCLIP: Efficient Cross-Modal Pre-training by Ensemble Confident Learning and Language Modeling 10 Sep 2021 · 0 repositories · arXiv:2109.04699
-
Improving Video-Text Retrieval by Multi-Stream Corpus Alignment and Dual Softmax Loss 9 Sep 2021 · 2 repositories · arXiv:2109.04290
-
Zero-Shot Out-of-Distribution Detection Based on the Pre-trained Model CLIP 6 Sep 2021 · 2 repositories · arXiv:2109.02748Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 10 unverified (of 11 harvested samples)
-
Robust fine-tuning of zero-shot models 4 Sep 2021 · 3 repositories · arXiv:2109.01903Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 3 harvested samples) · 1 pointer-only (licence)
-
Learning to Prompt for Vision-Language Models 2 Sep 2021 · 18 repositories · arXiv:2109.01134Syntology community repositories only · 3 ran (of which 1 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 1 where Syntology's instrument failed) · 9 unverified (of 12 harvested samples)
-
LIGAR: Lightweight General-purpose Action Recognition 30 Aug 2021 · 1 repository · arXiv:2108.13153
-
Contrastive Language-Image Pre-training for the Italian Language 19 Aug 2021 · 1 repository · arXiv:2108.08688
-
Evaluating CLIP: Towards Characterization of Broader Capabilities and Downstream Implications 5 Aug 2021 · 0 repositories · arXiv:2108.02818
-
Is Object Detection Necessary for Human-Object Interaction Recognition? 27 Jul 2021 · 0 repositories · arXiv:2107.13083
-
Segmentation in Style: Unsupervised Semantic Image Segmentation with Stylegan and CLIP 26 Jul 2021 · 1 repository · arXiv:2107.12518Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Learning Concept Lengths Accelerates Concept Learning in ALC 10 Jul 2021 · 1 repository · arXiv:2107.04911
-
Exploiting the relationship between visual and textual features in social networks for image classification with zero-shot deep learning 8 Jul 2021 · 0 repositories · arXiv:2107.03751
-
In-distribution adversarial attacks on object recognition models using gradient-free search 30 Jun 2021 · 2 repositories · arXiv:2106.16198Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
AudioCLIP: Extending CLIP to Image, Text and Audio 24 Jun 2021 · 4 repositories · arXiv:2106.13043Syntology community repositories only · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 6 harvested samples)
-
Poisoning and Backdooring Contrastive Learning 17 Jun 2021 · 1 repository · arXiv:2106.09667
-
A Fair and Comprehensive Comparison of Multimodal Tweet Sentiment Analysis Methods 16 Jun 2021 · 1 repository · arXiv:2106.08829
-
Partial success in closing the gap between human and machine vision 14 Jun 2021 · 1 repository · arXiv:2106.07411Syntology official: harvested, nothing ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples) · 1 pointer-only (licence)
-
Assessing Multilingual Fairness in Pre-trained Multimodal Representations 12 Jun 2021 · 0 repositories · arXiv:2106.06683
-
ImaginE: An Imagination-Based Automatic Evaluation Metric for Natural Language Generation 10 Jun 2021 · 0 repositories · arXiv:2106.05970
-
Exploring the Limits of Out-of-Distribution Detection 6 Jun 2021 · 1 repository · arXiv:2106.03004
-
CLIP: A Dataset for Extracting Action Items for Physicians from Hospital Discharge Notes 4 Jun 2021 · 1 repository · arXiv:2106.02524
-
Personalizing Pre-trained Models 2 Jun 2021 · 0 repositories · arXiv:2106.01499
-
CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval 18 Apr 2021 · 5 repositories · arXiv:2104.08860Syntology official: harvested, nothing ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples) · 3 pointer-only (licence)
-
Data-Efficient Language-Supervised Zero-Shot Learning with Self-Distillation 18 Apr 2021 · 0 repositories · arXiv:2104.08945
-
SI-Score: An image dataset for fine-grained analysis of robustness to object location, rotation and size 9 Apr 2021 · 1 repository · arXiv:2104.04191Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Putting NeRF on a Diet: Semantically Consistent Few-Shot View Synthesis 1 Apr 2021 · 2 repositories · arXiv:2104.00677Syntology official (archive's flag): 1 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 4 harvested samples) · 2 pointer-only (licence)
-
CLIP: Cheap Lipschitz Training of Neural Networks 23 Mar 2021 · 1 repository · arXiv:2103.12531Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 7 unverified (of 10 harvested samples)
-
Reading Isn't Believing: Adversarial Attacks On Multi-Modal Neurons 18 Mar 2021 · 0 repositories · arXiv:2103.10480
-
WenLan: Bridging Vision and Language by Large-Scale Multi-Modal Pre-Training 11 Mar 2021 · 2 repositories · arXiv:2103.06561
-
Learning Transferable Visual Models From Natural Language Supervision 26 Feb 2021 · 82 repositories · arXiv:2103.00020Syntology official: no sample here; runs from other or unrecorded repositories · 16 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 1 honoured, 1 violated, 0 with no contract checked; 14 where Syntology's instrument failed) · 4 unverified (of 20 harvested samples) · 16 pointer-only (licence)