Methods › Computer Vision › Vision and Language Pre-Trained Models › CLIP › Papers, page 28
Contrastive Language-Image Pre-training
CLIP
Papers archive 2025-07-28
archive papers tagged: 3,094 · with a code link: 1,617 · where Syntology ran a sample: 649 (554 with a run with no instrument failure, 95 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (649 of 3,094 tagged: 554 with a run with no instrument failure, 95 where every run was a failure of Syntology's instrument)
Page 28 of 31: papers 2,701 to 2,800 of 3,094, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
SuS-X: Training-Free Name-Only Transfer of Vision-Language Models 28 Nov 2022 · 2 repositories · arXiv:2211.16198Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 2 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
SegCLIP: Patch Aggregation with Learnable Centers for Open-Vocabulary Semantic Segmentation 27 Nov 2022 · 1 repository · arXiv:2211.14813
-
CLIP-ReID: Exploiting Vision-Language Model for Image Re-Identification without Concrete Text Labels 25 Nov 2022 · 2 repositories · arXiv:2211.13977
-
ComCLIP: Training-Free Compositional Image and Text Matching 25 Nov 2022 · 1 repository · arXiv:2211.13854
-
On the Importance of Image Encoding in Automated Chest X-Ray Report Generation 24 Nov 2022 · 1 repository · arXiv:2211.13465
-
Shifted Diffusion for Text-to-image Generation 24 Nov 2022 · 1 repository · arXiv:2211.15388Syntology official (archive's flag): 13 ran · 13 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 2 honoured, 1 violated, 7 with no contract checked; 3 where Syntology's instrument failed) · 4 unverified (of 17 harvested samples) · 6 pointer-only (licence)
-
Holistic Visual-Textual Sentiment Analysis with Prior Models 23 Nov 2022 · 1 repository · arXiv:2211.12981
-
InDiReCT: Language-Guided Zero-Shot Deep Metric Learning for Images 23 Nov 2022 · 1 repository · arXiv:2211.12760
-
Schrödinger's Bat: Diffusion Models Sometimes Generate Polysemous Words in Superposition 23 Nov 2022 · 1 repository · arXiv:2211.13095
-
Texts as Images in Prompt Tuning for Multi-Label Image Recognition 23 Nov 2022 · 1 repository · arXiv:2211.12739
-
VoP: Text-Video Co-operative Prompt Tuning for Cross-Modal Retrieval 23 Nov 2022 · 1 repository · arXiv:2211.12764
-
On the Transferability of Visual Features in Generalized Zero-Shot Learning 22 Nov 2022 · 1 repository · arXiv:2211.12494
-
Retrieval-Augmented Multimodal Language Modeling 22 Nov 2022 · 0 repositories · arXiv:2211.12561
-
Expectation-Maximization Contrastive Learning for Compact Video-and-Language Representations 21 Nov 2022 · 4 repositories · arXiv:2211.11427Syntology official (archive's flag): 1 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples)
-
Language in a Bottle: Language Model Guided Concept Bottlenecks for Interpretable Image Classification 21 Nov 2022 · 2 repositories · arXiv:2211.11158
-
LISA: Localized Image Stylization with Audio via Implicit Neural Representation 21 Nov 2022 · 0 repositories · arXiv:2211.11381
-
PointCLIP V2: Prompting CLIP and GPT for Powerful 3D Open-world Learning 21 Nov 2022 · 2 repositories · arXiv:2211.11682Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 1 violated, 3 with no contract checked; 4 where Syntology's instrument failed) · 4 unverified (of 12 harvested samples) · 4 pointer-only (licence)
-
Robotic Skill Acquisition via Instruction Augmentation with Vision-Language Models 21 Nov 2022 · 0 repositories · arXiv:2211.11736
-
Understanding and Improving Visual Prompting: A Label-Mapping Perspective 21 Nov 2022 · 1 repository · arXiv:2211.11635
-
IC3D: Image-Conditioned 3D Diffusion for Shape Generation 20 Nov 2022 · 0 repositories · arXiv:2211.10865
-
CAE v2: Context Autoencoder with CLIP Target 17 Nov 2022 · 0 repositories · arXiv:2211.09799
-
Cross-Modal Adapter for Text-Video Retrieval 17 Nov 2022 · 1 repository · arXiv:2211.09623Syntology official (archive's flag): 15 ran · 15 ran (of which 0 constructed an object rather than computing a result; 13 with no instrument failure: 1 honoured, 2 violated, 10 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 15 harvested samples) · 6 pointer-only (licence)
-
GLAMI-1M: A Multilingual Image-Text Fashion Dataset 17 Nov 2022 · 1 repository · arXiv:2211.14451
-
TempNet: Temporal Attention Towards the Detection of Animal Behaviour in Videos 17 Nov 2022 · 0 repositories · arXiv:2211.09950
-
Robust Online Video Instance Segmentation with Track Queries 16 Nov 2022 · 1 repository · arXiv:2211.09108
-
Federated Adaptive Prompt Tuning for Multi-Domain Collaborative Learning 15 Nov 2022 · 1 repository · arXiv:2211.07864Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
ContextCLIP: Contextual Alignment of Image-Text pairs on CLIP visual representations 14 Nov 2022 · 0 repositories · arXiv:2211.07122
-
EVA: Exploring the Limits of Masked Visual Representation Learning at Scale 14 Nov 2022 · 6 repositories · arXiv:2211.07636Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 3 harvested samples)
-
A Novel Sampling Scheme for Text- and Image-Conditional Image Synthesis in Quantized Latent Spaces 14 Nov 2022 · 4 repositories · arXiv:2211.07292Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 3 harvested samples)
-
Zero-shot Image Captioning by Anchor-augmented Vision-Language Space Alignment 14 Nov 2022 · 0 repositories · arXiv:2211.07275
-
AltCLIP: Altering the Language Encoder in CLIP for Extended Language Capabilities 12 Nov 2022 · 2 repositories · arXiv:2211.06679Syntology official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 11 harvested samples) · 2 pointer-only (licence)
-
Foundation Models for Semantic Novelty in Reinforcement Learning 9 Nov 2022 · 0 repositories · arXiv:2211.04878
-
Disentangling Content and Motion for Text-Based Neural Video Manipulation 5 Nov 2022 · 1 repository · arXiv:2211.02980
-
Domain Adaptive Video Semantic Segmentation via Cross-Domain Moving Object Mixing 4 Nov 2022 · 1 repository · arXiv:2211.02307
-
Understanding and Mitigating Overfitting in Prompt Tuning for Vision-Language Models 4 Nov 2022 · 1 repository · arXiv:2211.02219
-
Chinese CLIP: Contrastive Vision-Language Pretraining in Chinese 2 Nov 2022 · 1 repository · arXiv:2211.01335Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 2 honoured, 0 violated, 3 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 7 harvested samples) · 3 pointer-only (licence)
-
eDiff-I: Text-to-Image Diffusion Models with an Ensemble of Expert Denoisers 2 Nov 2022 · 2 repositories · arXiv:2211.01324Syntology 11 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 1 honoured, 0 violated, 9 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 13 harvested samples)
-
M-SpeechCLIP: Leveraging Large-Scale, Pre-Trained Models for Multilingual Speech to Image Retrieval 2 Nov 2022 · 0 repositories · arXiv:2211.01180
-
MuMIC -- Multimodal Embedding for Multi-label Image Classification with Tempered Sigmoid 2 Nov 2022 · 0 repositories · arXiv:2211.05232
-
CLIP-Sculptor: Zero-Shot Generation of High-Fidelity and Diverse Shapes from Natural Language 2 Nov 2022 · 0 repositories · arXiv:2211.01427
-
Text-Only Training for Image Captioning using Noise-Injected CLIP 1 Nov 2022 · 4 repositories · arXiv:2211.00575Syntology official (archive's flag): 2 ran · 4 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 1 honoured, 1 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples) · 1 pointer-only (licence)
-
VID-Trans-ReID: Enhanced Video Transformers for Person Re-identification 1 Nov 2022 · 1 repository
-
ImagineNET: Target Speaker Extraction with Intermittent Visual Cue through Embedding Inpainting 31 Oct 2022 · 1 repository · arXiv:2211.00109
-
Image-free Domain Generalization via CLIP for 3D Hand Pose Estimation 30 Oct 2022 · 0 repositories · arXiv:2210.16788
-
Unsupervised Audio-Visual Lecture Segmentation 29 Oct 2022 · 1 repository · arXiv:2210.16644
-
Line Spectral Estimation via Unlimited Sampling 28 Oct 2022 · 0 repositories · arXiv:2210.15811
-
Being Comes from Not-being: Open-vocabulary Text-to-Motion Generation with Wordless Training 28 Oct 2022 · 1 repository · arXiv:2210.15929Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Do Pre-trained Models Benefit Equally in Continual Learning? 27 Oct 2022 · 1 repository · arXiv:2210.15701
-
Towards Reliable Zero Shot Classification in Self-Supervised Models with Conformal Prediction 27 Oct 2022 · 0 repositories · arXiv:2210.15805
-
FairCLIP: Social Bias Elimination based on Attribute Prototype Learning and Representation Neutralization 26 Oct 2022 · 0 repositories · arXiv:2210.14562
-
Visual Answer Localization with Cross-modal Mutual Knowledge Transfer 26 Oct 2022 · 1 repository · arXiv:2210.14823
-
Inferring Past Human Actions in Homes with Abductive Reasoning 24 Oct 2022 · 1 repository · arXiv:2210.13984
-
Language-free Training for Zero-shot Video Grounding 24 Oct 2022 · 0 repositories · arXiv:2210.12977
-
The Robustness Limits of SoTA Vision Models to Natural Variation 24 Oct 2022 · 0 repositories · arXiv:2210.13604
-
BASQ: Branch-wise Activation-clipping Search Quantization for Sub-4-bit Neural Networks 23 Oct 2022 · 1 repository
-
Towards Real-Time Text2Video via CLIP-Guided, Pixel-Level Optimization 23 Oct 2022 · 1 repository · arXiv:2210.12826
-
3DALL-E: Integrating Text-to-Image AI in 3D Design Workflows 20 Oct 2022 · 0 repositories · arXiv:2210.11603
-
General Image Descriptors for Open World Image Retrieval using ViT CLIP 20 Oct 2022 · 1 repository · arXiv:2210.11141Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
MovieCLIP: Visual Scene Recognition in Movies 20 Oct 2022 · 1 repository · arXiv:2210.11065Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 1 honoured, 0 violated, 6 with no contract checked; 1 where Syntology's instrument failed) · 3 unverified (of 11 harvested samples) · 2 pointer-only (licence)
-
TANGO: Text-driven Photorealistic and Robust 3D Stylization via Lighting Decomposition 20 Oct 2022 · 1 repository · arXiv:2210.11277
-
CLIP-Driven Fine-grained Text-Image Person Re-identification 19 Oct 2022 · 1 repository · arXiv:2210.10276Syntology official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 2 honoured, 0 violated, 0 with no contract checked; 8 where Syntology's instrument failed) · 2 unverified (of 12 harvested samples) · 6 pointer-only (licence)
-
CPL: Counterfactual Prompt Learning for Vision and Language Models 19 Oct 2022 · 0 repositories · arXiv:2210.10362
-
5th Place Solution to Kaggle Google Universal Image Embedding Competition 18 Oct 2022 · 1 repository · arXiv:2210.09495
-
MedCLIP: Contrastive Learning from Unpaired Medical Images and Text 18 Oct 2022 · 1 repository · arXiv:2210.10163
-
Probing Cross-modal Semantics Alignment Capability from the Textual Perspective 18 Oct 2022 · 0 repositories · arXiv:2210.09550
-
6th Place Solution to Google Universal Image Embedding 17 Oct 2022 · 0 repositories · arXiv:2210.09377
-
Contrastive Language-Image Pre-Training with Knowledge Graphs 17 Oct 2022 · 0 repositories · arXiv:2210.08901
-
Non-Contrastive Learning Meets Language-Image Pre-Training 17 Oct 2022 · 1 repository · arXiv:2210.09304
-
Track Targets by Dense Spatio-Temporal Position Encoding 17 Oct 2022 · 0 repositories · arXiv:2210.09455
-
LAION-5B: An open large-scale dataset for training next generation image-text models 16 Oct 2022 · 5 repositories · arXiv:2210.08402Syntology official: no sample here; runs from other or unrecorded repositories · 14 ran (of which 0 constructed an object rather than computing a result; 12 with no instrument failure: 0 honoured, 1 violated, 11 with no contract checked; 2 where Syntology's instrument failed) · 4 unverified (of 18 harvested samples) · 3 pointer-only (licence)
-
One Model to Edit Them All: Free-Form Text-Driven Image Manipulation with Semantic Modulations 14 Oct 2022 · 1 repository · arXiv:2210.07883
-
Caption supervision enables robust learners 13 Oct 2022 · 1 repository · arXiv:2210.07396
-
Unified Vision and Language Prompt Learning 13 Oct 2022 · 1 repository · arXiv:2210.07225
-
Visual Classification via Description from Large Language Models 13 Oct 2022 · 3 repositories · arXiv:2210.07183Syntology official (archive's flag): 2 ran · 6 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 6 where Syntology's instrument failed) · 0 unverified (of 6 harvested samples) · 3 pointer-only (licence)
-
Hate-CLIPper: Multimodal Hateful Meme Classification based on Cross-modal Interaction of CLIP Features 12 Oct 2022 · 1 repository · arXiv:2210.05916Syntology official: harvested, nothing ran · 0 ran · 5 unverified (of 5 harvested samples)
-
CLIP also Understands Text: Prompting CLIP for Phrase Understanding 11 Oct 2022 · 0 repositories · arXiv:2210.05836
-
CLIP-Fields: Weakly Supervised Semantic Fields for Robotic Memory 11 Oct 2022 · 2 repositories · arXiv:2210.05663Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Transfer Learning with Joint Fine-Tuning for Multimodal Sentiment Analysis 11 Oct 2022 · 1 repository · arXiv:2210.05790
-
Unifying Diffusion Models' Latent Space, with Applications to CycleDiffusion and Guidance 11 Oct 2022 · 4 repositories · arXiv:2210.05559Syntology official (archive's flag): 1 ran · 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 1 honoured, 1 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Automated Audio Captioning via Fusion of Low- and High- Dimensional Features 10 Oct 2022 · 0 repositories · arXiv:2210.05037
-
Bridging CLIP and StyleGAN through Latent Alignment for Image Editing 10 Oct 2022 · 0 repositories · arXiv:2210.04506
-
ConTra: (Con)text (Tra)nsformer for Cross-Modal Video Retrieval 9 Oct 2022 · 1 repository · arXiv:2210.04341
-
Learning to Decompose Visual Features with Latent Textual Prompts 9 Oct 2022 · 0 repositories · arXiv:2210.04287
-
Open-Vocabulary Semantic Segmentation with Mask-adapted CLIP 9 Oct 2022 · 1 repository · arXiv:2210.04150Syntology official: harvested, nothing ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
CLIP-PAE: Projection-Augmentation Embedding to Extract Relevant Features for a Disentangled, Interpretable, and Controllable Text-Guided Face Manipulation 8 Oct 2022 · 0 repositories · arXiv:2210.03919
-
FastCLIPstyler: Optimisation-free Text-based Image Style Transfer Using Style Representations 7 Oct 2022 · 0 repositories · arXiv:2210.03461
-
SVL-Adapter: Self-Supervised Adapter for Vision-Language Pretrained Models 7 Oct 2022 · 1 repository · arXiv:2210.03794
-
CLIP model is an Efficient Continual Learner 6 Oct 2022 · 1 repository · arXiv:2210.03114Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
MaPLe: Multi-modal Prompt Learning 6 Oct 2022 · 3 repositories · arXiv:2210.03117Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 3 where Syntology's instrument failed) · 2 unverified (of 6 harvested samples) · 2 pointer-only (licence)
-
Real-World Robot Learning with Masked Visual Pre-training 6 Oct 2022 · 1 repository · arXiv:2210.03109
-
VLSNR:Vision-Linguistics Coordination Time Sequence-aware News Recommendation 6 Oct 2022 · 2 repositories · arXiv:2210.02946
-
clip2latent: Text driven sampling of a pre-trained StyleGAN using denoising diffusion and CLIP 5 Oct 2022 · 2 repositories · arXiv:2210.02347
-
Bayesian Prompt Learning for Image-Language Model Generalization 5 Oct 2022 · 1 repository · arXiv:2210.02390Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 3 where Syntology's instrument failed) · 2 unverified (of 6 harvested samples) · 3 pointer-only (licence)
-
When and why vision-language models behave like bags-of-words, and what to do about it? 4 Oct 2022 · 1 repository · arXiv:2210.01936Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 5 unverified (of 8 harvested samples) · 7 pointer-only (licence)
-
CLIP2Point: Transfer CLIP to Point Cloud Classification with Image-Depth Pre-training 3 Oct 2022 · 1 repository · arXiv:2210.01055
-
LASP: Text-to-Text Optimization for Language-Aware Soft Prompting of Vision & Language Models 3 Oct 2022 · 1 repository · arXiv:2210.01115Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 3 where Syntology's instrument failed) · 3 unverified (of 7 harvested samples) · 2 pointer-only (licence)
-
PLOT: Prompt Learning with Optimal Transport for Vision-Language Models 3 Oct 2022 · 1 repository · arXiv:2210.01253Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 4 harvested samples)
-
SpeechCLIP: Integrating Speech with Pre-Trained Vision and Language Model 3 Oct 2022 · 1 repository · arXiv:2210.00705
-
Improving ProtoNet for Few-Shot Video Object Recognition: Winner of ORBIT Challenge 2022 1 Oct 2022 · 4 repositories · arXiv:2210.00174
-
Data Poisoning Attacks Against Multimodal Encoders 30 Sep 2022 · 1 repository · arXiv:2209.15266Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)