Methods › Computer Vision › Vision and Language Pre-Trained Models › CLIP › Papers, page 24
Contrastive Language-Image Pre-training
CLIP
Papers archive 2025-07-28
archive papers tagged: 3,094 · with a code link: 1,617 · where Syntology ran a sample: 649 (554 with a run with no instrument failure, 95 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (649 of 3,094 tagged: 554 with a run with no instrument failure, 95 where every run was a failure of Syntology's instrument)
Page 24 of 31: papers 2,301 to 2,400 of 3,094, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Retrieval-Enhanced Visual Prompt Learning for Few-shot Classification 4 Jun 2023 · 0 repositories · arXiv:2306.02243
-
Concurrent Classifier Error Detection (CCED) in Large Scale Machine Learning Systems 2 Jun 2023 · 0 repositories · arXiv:2306.01820
-
Enhancing CLIP with CLIP: Exploring Pseudolabeling for Limited-Label Prompt Tuning 2 Jun 2023 · 2 repositories · arXiv:2306.01669
-
LoCoOp: Few-Shot Out-of-Distribution Detection via Prompt Learning 2 Jun 2023 · 2 repositories · arXiv:2306.01293Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 1 honoured, 0 violated, 4 with no contract checked; 4 where Syntology's instrument failed) · 3 unverified (of 12 harvested samples) · 7 pointer-only (licence)
-
Exploring the Versatility of Zero-Shot CLIP for Interstitial Lung Disease Classification 1 Jun 2023 · 0 repositories · arXiv:2306.01111
-
Discovering Failure Modes of Text-guided Diffusion Models via Adversarial Search 1 Jun 2023 · 0 repositories · arXiv:2306.00974
-
SnapFusion: Text-to-Image Diffusion Model on Mobile Devices within Two Seconds 1 Jun 2023 · 0 repositories · arXiv:2306.00980
-
StableRep: Synthetic Images from Text-to-Image Models Make Strong Visual Representation Learners 1 Jun 2023 · 2 repositories · arXiv:2306.00984Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
UniDiff: Advancing Vision-Language Models with Generative and Discriminative Learning 1 Jun 2023 · 0 repositories · arXiv:2306.00813
-
Dense and Aligned Captions (DAC) Promote Compositional Reasoning in VL Models 31 May 2023 · 1 repository · arXiv:2305.19595
-
Improving CLIP Training with Language Rewrites 31 May 2023 · 1 repository · arXiv:2305.20088Syntology official (archive's flag): 2 ran · 2 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples)
-
Label-Retrieval-Augmented Diffusion Models for Learning from Noisy Labels 31 May 2023 · 1 repository · arXiv:2305.19518Syntology official (archive's flag): 7 ran · 7 ran (of which 4 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 3 where Syntology's instrument failed) · 2 unverified (of 9 harvested samples)
-
LMCap: Few-shot Multilingual Image Captioning by Retrieval Augmented Language Model Prompting 31 May 2023 · 1 repository · arXiv:2305.19821
-
Using Visual Cropping to Enhance Fine-Detail Question Answering of BLIP-Family Models 31 May 2023 · 0 repositories · arXiv:2306.00228
-
DisCLIP: Open-Vocabulary Referring Expression Generation 30 May 2023 · 0 repositories · arXiv:2305.19108
-
Scalable Performance Analysis for Vision-Language Models 30 May 2023 · 1 repository · arXiv:2305.18786
-
Deeply Coupled Cross-Modal Prompt Learning 29 May 2023 · 1 repository · arXiv:2305.17903
-
Federated Learning of Gboard Language Models with Differential Privacy 29 May 2023 · 1 repository · arXiv:2305.18465
-
GlyphControl: Glyph Conditional Control for Visual Text Generation 29 May 2023 · 1 repository · arXiv:2305.18259Syntology official (archive's flag): 4 ran · 4 ran (of which 3 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 5 unverified (of 9 harvested samples)
-
Reconstructing the Mind's Eye: fMRI-to-Image with Contrastive Learning and Diffusion Priors 29 May 2023 · 1 repository · arXiv:2305.18274Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 7 harvested samples)
-
Test-Time Adaptation with CLIP Reward for Zero-Shot Generalization in Vision-Language Models 29 May 2023 · 1 repository · arXiv:2305.18010Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 2 where Syntology's instrument failed) · 5 unverified (of 10 harvested samples) · 2 pointer-only (licence)
-
The Rise of AI Language Pathologists: Exploring Two-level Prompt Learning for Few-shot Weakly-supervised Whole Slide Image Classification 29 May 2023 · 1 repository · arXiv:2305.17891Syntology official (archive's flag): 4 ran · 4 ran (of which 2 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 1 violated, 2 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
APRIL-GAN: A Zero-/Few-Shot Anomaly Classification and Segmentation Method for CVPR 2023 VAND Workshop Challenge Tracks 1&2: 1st Place on Zero-shot AD and 4th Place on Few-shot AD 27 May 2023 · 2 repositories · arXiv:2305.17382Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 1 violated, 3 with no contract checked; 2 where Syntology's instrument failed) · 4 unverified (of 10 harvested samples) · 7 pointer-only (licence)
-
CAILA: Concept-Aware Intra-Layer Adapters for Compositional Zero-Shot Learning 26 May 2023 · 2 repositories · arXiv:2305.16681
-
GeoVLN: Learning Geometry-Enhanced Visual Representation with Slot Attention for Vision-and-Language Navigation 26 May 2023 · 1 repository · arXiv:2305.17102
-
Learning to Imagine: Visually-Augmented Natural Language Generation 26 May 2023 · 1 repository · arXiv:2305.16944Syntology official: no sample here; runs from other or unrecorded repositories · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
On Evaluating Adversarial Robustness of Large Vision-Language Models 26 May 2023 · 1 repository · arXiv:2305.16934Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 1 honoured, 0 violated, 1 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 2 pointer-only (licence)
-
OpenVIS: Open-vocabulary Video Instance Segmentation 26 May 2023 · 1 repository · arXiv:2305.16835
-
Are Diffusion Models Vision-And-Language Reasoners? 25 May 2023 · 1 repository · arXiv:2305.16397Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
DiffCLIP: Leveraging Stable Diffusion for Language Grounded 3D Classification 25 May 2023 · 0 repositories · arXiv:2305.15957
-
Text-to-Motion Retrieval: Towards Joint Understanding of Human Motion Data and Natural Language 25 May 2023 · 1 repository · arXiv:2305.15842
-
Uni-ControlNet: All-in-One Control to Text-to-Image Diffusion Models 25 May 2023 · 1 repository · arXiv:2305.16322Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 2 violated, 3 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 8 harvested samples) · 8 pointer-only (licence)
-
ChatFace: Chat-Guided Real Face Editing via Diffusion Latent Space Manipulation 24 May 2023 · 0 repositories · arXiv:2305.14742
-
Decomposing Complex Queries for Tip-of-the-tongue Retrieval 24 May 2023 · 0 repositories · arXiv:2305.15053
-
PathAsst: A Generative Foundation AI Assistant Towards Artificial General Intelligence of Pathology 24 May 2023 · 1 repository · arXiv:2305.15072
-
Alt-Text with Context: Improving Accessibility for Images on Twitter 24 May 2023 · 0 repositories · arXiv:2305.14779
-
Text encoders bottleneck compositionality in contrastive vision-language models 24 May 2023 · 1 repository · arXiv:2305.14897
-
Can Language Models Understand Physical Concepts? 23 May 2023 · 1 repository · arXiv:2305.14057
-
CLIP4STR: A Simple Baseline for Scene Text Recognition with Pre-trained Vision-Language Model 23 May 2023 · 1 repository · arXiv:2305.14014Syntology 6 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 3 where Syntology's instrument failed) · 4 unverified (of 10 harvested samples) · 2 pointer-only (licence)
-
Coarse-to-Fine Contrastive Learning in Image-Text-Graph Space for Improved Vision-Language Compositionality 23 May 2023 · 0 repositories · arXiv:2305.13812
-
CPNet: Exploiting CLIP-based Attention Condenser and Probability Map Guidance for High-fidelity Talking Face Generation 23 May 2023 · 0 repositories · arXiv:2305.13962
-
Cross3DVG: Cross-Dataset 3D Visual Grounding on Different RGB-D Scans 23 May 2023 · 1 repository · arXiv:2305.13876
-
Parts of Speech-Grounded Subspaces in Vision-Language Models 23 May 2023 · 2 repositories · arXiv:2305.14053Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 4 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
Prompting Language-Informed Distribution for Compositional Zero-Shot Learning 23 May 2023 · 1 repository · arXiv:2305.14428Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 3 where Syntology's instrument failed) · 1 unverified (of 8 harvested samples) · 3 pointer-only (licence)
-
S-CLIP: Semi-supervised Vision-Language Learning using Few Specialist Captions 23 May 2023 · 1 repository · arXiv:2305.14095Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Training Transitive and Commutative Multimodal Transformers with LoReTTa 23 May 2023 · 0 repositories · arXiv:2305.14243
-
Weakly Supervised 3D Open-vocabulary Segmentation 23 May 2023 · 1 repository · arXiv:2305.14093Syntology official (archive's flag): 12 ran · 12 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 2 where Syntology's instrument failed) · 2 unverified (of 14 harvested samples) · 10 pointer-only (licence)
-
Connecting Multi-modal Contrastive Representations 22 May 2023 · 0 repositories · arXiv:2305.14381
-
ControlVideo: Training-free Controllable Text-to-Video Generation 22 May 2023 · 1 repository · arXiv:2305.13077Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 1 where Syntology's instrument failed) · 6 unverified (of 9 harvested samples) · 4 pointer-only (licence)
-
LaDI-VTON: Latent Diffusion Textual-Inversion Enhanced Virtual Try-On 22 May 2023 · 1 repository · arXiv:2305.13501
-
The CLIP Model is Secretly an Image-to-Prompt Converter 22 May 2023 · 0 repositories · arXiv:2305.12716
-
Towards Explainable In-the-Wild Video Quality Assessment: A Database and a Language-Prompted Approach 22 May 2023 · 1 repository · arXiv:2305.12726
-
VLAB: Enhancing Video Language Pre-training by Feature Adapting and Blending 22 May 2023 · 0 repositories · arXiv:2305.13167
-
Your smartphone could act as a pulse-oximeter and as a single-lead ECG 21 May 2023 · 0 repositories · arXiv:2305.12583
-
Zero-shot Visual Relation Detection via Composite Visual Cues from Large Language Models 21 May 2023 · 1 repository · arXiv:2305.12476Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Boosting Human-Object Interaction Detection with Text-to-Image Diffusion Model 20 May 2023 · 1 repository · arXiv:2305.12252
-
Movie101: A New Movie Understanding Benchmark 20 May 2023 · 1 repository · arXiv:2305.12140
-
What Makes for Good Visual Tokenizers for Large Language Models? 20 May 2023 · 1 repository · arXiv:2305.12223Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 1 pointer-only (licence)
-
AttriCLIP: A Non-Incremental Learner for Incremental Knowledge Learning 19 May 2023 · 1 repository · arXiv:2305.11488
-
CM-MaskSD: Cross-Modality Masked Self-Distillation for Referring Image Segmentation 19 May 2023 · 0 repositories · arXiv:2305.11481
-
Efficient Cross-Lingual Transfer for Chinese Stable Diffusion with Images as Pivots 19 May 2023 · 0 repositories · arXiv:2305.11540
-
Federated Foundation Models: Privacy-Preserving and Collaborative Learning for Large Models 19 May 2023 · 0 repositories · arXiv:2305.11414
-
Instruct2Act: Mapping Multi-modality Instructions to Robotic Actions with Large Language Model 18 May 2023 · 1 repository · arXiv:2305.11176
-
LLMScore: Unveiling the Power of Large Language Models in Text-to-Image Synthesis Evaluation 18 May 2023 · 1 repository · arXiv:2305.11116
-
OpenShape: Scaling Up 3D Shape Representation Towards Open-World Understanding 18 May 2023 · 1 repository · arXiv:2305.10764Syntology 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 7 harvested samples) · 2 pointer-only (licence)
-
Universal Domain Adaptation from Foundation Models: A Baseline Study 18 May 2023 · 1 repository · arXiv:2305.11092
-
CLIP-GCD: Simple Language Guided Generalized Category Discovery 17 May 2023 · 0 repositories · arXiv:2305.10420
-
CLIP-VG: Self-paced Curriculum Adapting of CLIP for Visual Grounding 15 May 2023 · 3 repositories · arXiv:2305.08685
-
Improved baselines for vision-language pre-training 15 May 2023 · 1 repository · arXiv:2305.08675
-
Laughing Matters: Introducing Laughing-Face Generation using Diffusion Models 15 May 2023 · 1 repository · arXiv:2305.08854
-
CLIP-Count: Towards Text-Guided Zero-Shot Object Counting 12 May 2023 · 1 repository · arXiv:2305.07304Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 1 honoured, 0 violated, 2 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
An Inverse Scaling Law for CLIP Training 11 May 2023 · 1 repository · arXiv:2305.07017Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 11 harvested samples) · 2 pointer-only (licence)
-
Continual Vision-Language Representation Learning with Off-Diagonal Information 11 May 2023 · 0 repositories · arXiv:2305.07437
-
Learning the Visualness of Text Using Large Vision-Language Models 11 May 2023 · 0 repositories · arXiv:2305.10434
-
iEdit: Localised Text-guided Image Editing with Weak Supervision 10 May 2023 · 0 repositories · arXiv:2305.05947
-
Text-To-Concept (and Back) via Cross-Model Alignment 10 May 2023 · 1 repository · arXiv:2305.06386
-
A Review of Vision-Language Models and their Performance on the Hateful Memes Challenge 9 May 2023 · 1 repository · arXiv:2305.06159
-
Boosting Visual-Language Models by Exploiting Hard Samples 9 May 2023 · 1 repository · arXiv:2305.05208
-
Less is More: Removing Text-regions Improves CLIP Training Efficiency and Robustness 8 May 2023 · 1 repository · arXiv:2305.05095
-
LMPT: Prompt Tuning with Class-Specific Embedding Loss for Long-tailed Multi-Label Visual Recognition 8 May 2023 · 1 repository · arXiv:2305.04536Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 5 unverified (of 9 harvested samples)
-
Pick your Poison: Undetectability versus Robustness in Data Poisoning Attacks 7 May 2023 · 0 repositories · arXiv:2305.09671
-
High-fidelity Generalized Emotional Talking Face Generation with Multi-modal Emotion Space Learning 4 May 2023 · 0 repositories · arXiv:2305.02572
-
LLM2Loss: Leveraging Language Models for Explainable Model Diagnostics 4 May 2023 · 0 repositories · arXiv:2305.03212
-
Multimodal-driven Talking Face Generation via a Unified Diffusion-based Generator 4 May 2023 · 0 repositories · arXiv:2305.02594
-
Visual Transformation Telling 3 May 2023 · 1 repository · arXiv:2305.01928
-
Parameter-Efficient Cross-lingual Transfer of Vision and Language Models via Translation-based Alignment 2 May 2023 · 1 repository · arXiv:2305.03510Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 5 harvested samples)
-
CLIP-S⁴: Language-Guided Self-Supervised Semantic Segmentation 1 May 2023 · 0 repositories · arXiv:2305.01040
-
SceneGenie: Scene Graph Guided Diffusion Models for Image Synthesis 28 Apr 2023 · 0 repositories · arXiv:2304.14573
-
DataComp: In search of the next generation of multimodal datasets 27 Apr 2023 · 3 repositories · arXiv:2304.14108Syntology community repositories only · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Edit Everything: A Text-Guided Generative System for Images Editing 27 Apr 2023 · 1 repository · arXiv:2304.14006Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
IconShop: Text-Guided Vector Icon Synthesis with Autoregressive Transformers 27 Apr 2023 · 0 repositories · arXiv:2304.14400
-
From Association to Generation: Text-only Captioning by Unsupervised Cross-modal Mapping 26 Apr 2023 · 1 repository · arXiv:2304.13273
-
Training Large Scale Polynomial CNNs for E2E Inference over Homomorphic Encryption 26 Apr 2023 · 0 repositories · arXiv:2304.14836
-
TextDeformer: Geometry Manipulation using Text Guidance 26 Apr 2023 · 1 repository · arXiv:2304.13348Syntology official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 1 where Syntology's instrument failed) · 5 unverified (of 16 harvested samples) · 1 pointer-only (licence)
-
TR0N: Translator Networks for 0-Shot Plug-and-Play Conditional Generation 26 Apr 2023 · 2 repositories · arXiv:2304.13742Syntology official (archive's flag): 1 ran · 17 ran (of which 1 constructed an object rather than computing a result; 12 with no instrument failure: 0 honoured, 0 violated, 12 with no contract checked; 5 where Syntology's instrument failed) · 2 unverified (of 19 harvested samples) · 9 pointer-only (licence)
-
OFAR: A Multimodal Evidence Retrieval Framework for Illegal Live-streaming Identification 25 Apr 2023 · 0 repositories · arXiv:2304.12608
-
Stable and low-precision training for large-scale vision-language models 25 Apr 2023 · 1 repository · arXiv:2304.13013Syntology official: harvested for another paper · 0 ran · 3 unverified (of 3 harvested samples)
-
USA-Net: Unified Semantic and Affordance Representations for Robot Memory 24 Apr 2023 · 0 repositories · arXiv:2304.12164
-
Contrastive Language, Action, and State Pre-training for Robot Learning 21 Apr 2023 · 0 repositories · arXiv:2304.10782
-
RPLKG: Robust Prompt Learning with Knowledge Graph 21 Apr 2023 · 0 repositories · arXiv:2304.10805