Methods › Computer Vision › Vision and Language Pre-Trained Models › CLIP › Papers, page 27
Contrastive Language-Image Pre-training
CLIP
Papers archive 2025-07-28
archive papers tagged: 3,094 · with a code link: 1,617 · where Syntology ran a sample: 649 (554 with a run with no instrument failure, 95 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (649 of 3,094 tagged: 554 with a run with no instrument failure, 95 where every run was a failure of Syntology's instrument)
Page 27 of 31: papers 2,601 to 2,700 of 3,094, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Vision Learners Meet Web Image-Text Pairs 17 Jan 2023 · 0 repositories · arXiv:2301.07088
-
Multimodality Helps Unimodality: Cross-Modal Few-Shot Learning with Multimodal Models 16 Jan 2023 · 1 repository · arXiv:2301.06267Syntology official (archive's flag): 1 ran · 1 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified; the one sample that ran constructed an object rather than computing a result (of 1 harvested sample)
-
UATVR: Uncertainty-Adaptive Text-Video Retrieval 16 Jan 2023 · 1 repository · arXiv:2301.06309
-
It's Just a Matter of Time: Detecting Depression with Time-Enriched Multimodal Transformers 13 Jan 2023 · 1 repository · arXiv:2301.05453
-
CLIP2Scene: Towards Label-efficient 3D Scene Understanding by CLIP 12 Jan 2023 · 1 repository · arXiv:2301.04926
-
Logically at Factify 2: A Multi-Modal Fact Checking System Based on Evidence Retrieval techniques and Transformer Encoder Architecture 9 Jan 2023 · 0 repositories · arXiv:2301.03127
-
EgoDistill: Egocentric Head Motion Distillation for Efficient Video Understanding 5 Jan 2023 · 0 repositories · arXiv:2301.02217
-
FICE: Text-Conditioned Fashion Image Editing With Guided GAN Inversion 5 Jan 2023 · 1 repository · arXiv:2301.02110
-
StyleTalk: One-shot Talking Head Generation with Controllable Speaking Styles 3 Jan 2023 · 1 repository · arXiv:2301.01081Syntology official: harvested, nothing ran · 0 ran · 2 unverified (of 2 harvested samples)
-
CLIP-Driven Universal Model for Organ Segmentation and Tumor Detection 2 Jan 2023 · 2 repositories · arXiv:2301.00785Syntology community repositories only · 1 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified; the one sample that ran constructed an object rather than computing a result (of 2 harvested samples)
-
Muse: Text-To-Image Generation via Masked Generative Transformers 2 Jan 2023 · 5 repositories · arXiv:2301.00704Syntology 19 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 2 honoured, 5 violated, 1 with no contract checked; 11 where Syntology's instrument failed) · 2 unverified (of 21 harvested samples) · 11 pointer-only (licence)
-
A Simple Framework for Text-Supervised Semantic Segmentation 1 Jan 2023 · 1 repository
-
AutoAD II: The Sequel - Who, When, and What in Movie Audio Description 1 Jan 2023 · 0 repositories
-
CLIP-Cluster: CLIP-Guided Attribute Hallucination for Face Clustering 1 Jan 2023 · 0 repositories
-
CLIP-S4: Language-Guided Self-Supervised Semantic Segmentation 1 Jan 2023 · 0 repositories
-
CLIPPING: Distilling CLIP-Based Models With a Student Base for Video-Language Retrieval 1 Jan 2023 · 0 repositories
-
DIME-FM : DIstilling Multimodal and Efficient Foundation Models 1 Jan 2023 · 0 repositories
-
Exploring Intra-Class Variation Factors With Learnable Cluster Prompts for Semi-Supervised Image Synthesis 1 Jan 2023 · 0 repositories
-
Exploring Open-Vocabulary Semantic Segmentation from CLIP Vision Encoder Distillation Only 1 Jan 2023 · 0 repositories
-
Fusing Pre-Trained Language Models With Multimodal Prompts Through Reinforcement Learning 1 Jan 2023 · 1 repository
-
iCLIP: Bridging Image Classification and Contrastive Language-Image Pre-Training for Visual Recognition 1 Jan 2023 · 0 repositories
-
Improving CLIP Fine-tuning Performance 1 Jan 2023 · 1 repository
-
Learning Multi-Modal Class-Specific Tokens for Weakly Supervised Dense Object Localization 1 Jan 2023 · 0 repositories
-
LexLIP: Lexicon-Bottlenecked Language-Image Pre-Training for Large-Scale Image-Text Sparse Retrieval 1 Jan 2023 · 1 repository
-
MasQCLIP for Open-Vocabulary Universal Image Segmentation 1 Jan 2023 · 1 repository
-
Open-Set Fine-Grained Retrieval via Prompting Vision-Language Evaluator 1 Jan 2023 · 0 repositories
-
Ordered Atomic Activity for Fine-grained Interactive Traffic Scenario Understanding 1 Jan 2023 · 0 repositories
-
PADCLIP: Pseudo-labeling with Adaptive Debiasing in CLIP for Unsupervised Domain Adaptation 1 Jan 2023 · 0 repositories
-
PIDRo: Parallel Isomeric Attention with Dynamic Routing for Text-Video Retrieval 1 Jan 2023 · 0 repositories
-
RA-CLIP: Retrieval Augmented Contrastive Language-Image Pre-Training 1 Jan 2023 · 0 repositories
-
ReGen: A good Generative Zero-Shot Video Classifier Should be Rewarded 1 Jan 2023 · 0 repositories
-
Space-time Prompting for Video Class-incremental Learning 1 Jan 2023 · 0 repositories
-
Text-Guided Unsupervised Latent Transformation for Multi-Attribute Image Manipulation 1 Jan 2023 · 0 repositories
-
Tracking by Natural Language Specification with Long Short-term Context Decoupling 1 Jan 2023 · 0 repositories
-
Bidirectional Cross-Modal Knowledge Exploration for Video Recognition with Pre-trained Vision-Language Models 31 Dec 2022 · 5 repositories · arXiv:2301.00182
-
Cap4Video: What Can Auxiliary Captions Do for Text-Video Retrieval? 31 Dec 2022 · 4 repositories · arXiv:2301.00184Syntology official (archive's flag): 14 ran · 20 ran (of which 0 constructed an object rather than computing a result; 12 with no instrument failure: 2 honoured, 1 violated, 9 with no contract checked; 8 where Syntology's instrument failed) · 5 unverified (of 25 harvested samples) · 11 pointer-only (licence)
-
Unlearnable Clusters: Towards Label-agnostic Unlearnable Examples 31 Dec 2022 · 1 repository · arXiv:2301.01217
-
When are Lemons Purple? The Concept Association Bias of Vision-Language Models 22 Dec 2022 · 0 repositories · arXiv:2212.12043
-
3D Highlighter: Localizing Regions on 3D Shapes via Text Descriptions 21 Dec 2022 · 1 repository · arXiv:2212.11263
-
Contrastive Language-Vision AI Models Pretrained on Web-Scraped Multimodal Data Exhibit Sexual Objectification Bias 21 Dec 2022 · 1 repository · arXiv:2212.11261
-
Does CLIP Bind Concepts? Probing Compositionality in Large Image Models 20 Dec 2022 · 1 repository · arXiv:2212.10537
-
Tracking by Associating Clips 20 Dec 2022 · 0 repositories · arXiv:2212.10149
-
Unleashing the Power of Visual Prompting At the Pixel Level 20 Dec 2022 · 1 repository · arXiv:2212.10556Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples) · 1 pointer-only (licence)
-
Efficient Image Captioning for Edge Devices 18 Dec 2022 · 0 repositories · arXiv:2212.08985
-
3D Point Cloud Pre-training with Knowledge Distillation from 2D Images 17 Dec 2022 · 0 repositories · arXiv:2212.08974
-
Attentive Mask CLIP 16 Dec 2022 · 1 repository · arXiv:2212.08653
-
CLIP is Also an Efficient Segmenter: A Text-Driven Approach for Weakly Supervised Semantic Segmentation 16 Dec 2022 · 1 repository · arXiv:2212.09506Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
CLIPPO: Image-and-Language Understanding from Pixels Only 15 Dec 2022 · 1 repository · arXiv:2212.08045
-
MM-SHAP: A Performance-agnostic Metric for Measuring Multimodal Contributions in Vision and Language Models & Tasks 15 Dec 2022 · 1 repository · arXiv:2212.08158Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 5 unverified (of 10 harvested samples)
-
Visually-augmented pretrained language models for NLP tasks without images 15 Dec 2022 · 1 repository · arXiv:2212.07937
-
CLIPSep: Learning Text-queried Sound Separation with Noisy Unlabeled Videos 14 Dec 2022 · 1 repository · arXiv:2212.07065Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples)
-
EgoLoc: Revisiting 3D Object Localization from Egocentric Videos with Visual Queries 14 Dec 2022 · 1 repository · arXiv:2212.06969
-
NLIP: Noise-robust Language-Image Pre-training 14 Dec 2022 · 0 repositories · arXiv:2212.07086
-
Reproducible scaling laws for contrastive language-image learning 14 Dec 2022 · 5 repositories · arXiv:2212.07143Syntology community repositories only · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Significantly improving zero-shot X-ray pathology classification via fine-tuning pre-trained image-text encoders 14 Dec 2022 · 0 repositories · arXiv:2212.07050
-
Understanding Zero-Shot Adversarial Robustness for Large-Scale Models 14 Dec 2022 · 2 repositories · arXiv:2212.07016
-
LidarCLIP or: How I Learned to Talk to Point Clouds 13 Dec 2022 · 1 repository · arXiv:2212.06858
-
Localized Latent Updates for Fine-Tuning Vision-Language Models 13 Dec 2022 · 0 repositories · arXiv:2212.06556
-
On the Evolution of (Hateful) Memes by Means of Multimodal Contrastive Learning 13 Dec 2022 · 2 repositories · arXiv:2212.06573
-
CLIP Itself is a Strong Fine-tuner: Achieving 85.7% and 88.0% Top-1 Accuracy with ViT-B and ViT-L on ImageNet 12 Dec 2022 · 1 repository · arXiv:2212.06138
-
Doubly Right Object Recognition: A Why Prompt for Visual Rationales 12 Dec 2022 · 1 repository · arXiv:2212.06202
-
Reconstructing Humpty Dumpty: Multi-feature Graph Autoencoder for Open Set Action Recognition 12 Dec 2022 · 0 repositories · arXiv:2212.06023
-
CLIP-TSA: CLIP-Assisted Temporal Self-Attention for Weakly-Supervised Video Anomaly Detection 9 Dec 2022 · 2 repositories · arXiv:2212.05136
-
Multimodal Prototype-Enhanced Network for Few-Shot Action Recognition 9 Dec 2022 · 0 repositories · arXiv:2212.04873
-
Open Vocabulary Semantic Segmentation with Patch Aligned Contrastive Learning 9 Dec 2022 · 1 repository · arXiv:2212.04994
-
VindLU: A Recipe for Effective Video-and-Language Pretraining 9 Dec 2022 · 1 repository · arXiv:2212.05051
-
DialogCC: An Automated Pipeline for Creating High-Quality Multi-Modal Dialogue Dataset 8 Dec 2022 · 1 repository · arXiv:2212.04119
-
Diffusion Guided Domain Adaptation of Image Generators 8 Dec 2022 · 0 repositories · arXiv:2212.04473
-
EPCL: Frozen CLIP Transformer is An Efficient Point Cloud Encoder 8 Dec 2022 · 2 repositories · arXiv:2212.04098
-
Learning Domain Invariant Prompt for Vision-Language Models 8 Dec 2022 · 1 repository · arXiv:2212.04196
-
Learning to Dub Movies via Hierarchical Prosody Models 8 Dec 2022 · 1 repository · arXiv:2212.04054
-
Task Bias in Vision-Language Models 8 Dec 2022 · 0 repositories · arXiv:2212.04412
-
VASR: Visual Analogies of Situation Recognition 8 Dec 2022 · 1 repository · arXiv:2212.04542Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 8 harvested samples)
-
ZegCLIP: Towards Adapting CLIP for Zero-shot Semantic Segmentation 7 Dec 2022 · 1 repository · arXiv:2212.03588Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 1 honoured, 0 violated, 5 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 9 harvested samples) · 3 pointer-only (licence)
-
Adaptive Testing of Computer Vision Models 6 Dec 2022 · 1 repository · arXiv:2212.02774Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
Fine-tuned CLIP Models are Efficient Video Learners 6 Dec 2022 · 1 repository · arXiv:2212.03640Syntology official (archive's flag): 13 ran · 13 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 4 where Syntology's instrument failed) · 4 unverified (of 17 harvested samples) · 5 pointer-only (licence)
-
RANA: Relightable Articulated Neural Avatars 6 Dec 2022 · 0 repositories · arXiv:2212.03237
-
3D-LatentMapper: View Agnostic Single-View Reconstruction of 3D Shapes 5 Dec 2022 · 0 repositories · arXiv:2212.02184
-
CLIPVG: Text-Guided Image Manipulation Using Differentiable Vector Graphics 5 Dec 2022 · 1 repository · arXiv:2212.02122
-
Hierarchical Contrast for Unsupervised Skeleton-based Action Representation Learning 5 Dec 2022 · 1 repository · arXiv:2212.02082
-
Location-Aware Self-Supervised Transformers for Semantic Segmentation 5 Dec 2022 · 1 repository · arXiv:2212.02400Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
One-shot Implicit Animatable Avatars with Model-based Priors 5 Dec 2022 · 0 repositories · arXiv:2212.02469
-
Towards Generating Diverse Audio Captions via Adversarial Training 5 Dec 2022 · 0 repositories · arXiv:2212.02033
-
Improving Zero-shot Generalization and Robustness of Multi-modal Models 4 Dec 2022 · 1 repository · arXiv:2212.01758
-
CLIP: Train Faster with Less Data 2 Dec 2022 · 0 repositories · arXiv:2212.01452
-
ClipFace: Text-guided Editing of Textured 3D Morphable Models 2 Dec 2022 · 1 repository · arXiv:2212.01406
-
3D-LDM: Neural Implicit 3D Shape Generation with Latent Diffusion Models 1 Dec 2022 · 0 repositories · arXiv:2212.00842
-
Finetune like you pretrain: Improved finetuning of zero-shot vision models 1 Dec 2022 · 1 repository · arXiv:2212.00638Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 2 unverified (of 5 harvested samples) · 4 pointer-only (licence)
-
Focus! Relevant and Sufficient Context Selection for News Image Captioning 1 Dec 2022 · 0 repositories · arXiv:2212.00843
-
Improving Zero-Shot Models with Label Distribution Priors 1 Dec 2022 · 1 repository · arXiv:2212.00784
-
One-shot recognition of any material anywhere using contrastive learning with physics-based rendering 1 Dec 2022 · 1 repository · arXiv:2212.00648
-
Scaling Language-Image Pre-training via Masking 1 Dec 2022 · 6 repositories · arXiv:2212.00794Syntology community repositories only · 12 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 2 honoured, 5 violated, 2 with no contract checked; 3 where Syntology's instrument failed) · 2 unverified (of 14 harvested samples) · 4 pointer-only (licence)
-
CLIP-Nav: Using CLIP for Zero-Shot Vision-and-Language Navigation 30 Nov 2022 · 0 repositories · arXiv:2211.16649
-
Spatio-Temporal Crop Aggregation for Video Representation Learning 30 Nov 2022 · 0 repositories · arXiv:2211.17042
-
Context-Aware Robust Fine-Tuning 29 Nov 2022 · 0 repositories · arXiv:2211.16175
-
DATID-3D: Diversity-Preserved Domain Adaptation Using Text-to-Image Diffusion for 3D Generative Model 29 Nov 2022 · 0 repositories · arXiv:2211.16374
-
SinDDM: A Single Image Denoising Diffusion Model 29 Nov 2022 · 1 repository · arXiv:2211.16582Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 3 where Syntology's instrument failed) · 3 unverified (of 12 harvested samples) · 2 pointer-only (licence)
-
CLIP2GAN: Towards Bridging Text with the Latent Space of GANs 28 Nov 2022 · 0 repositories · arXiv:2211.15045
-
OpenScene: 3D Scene Understanding with Open Vocabularies 28 Nov 2022 · 1 repository · arXiv:2211.15654Syntology 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 9 harvested samples) · 8 pointer-only (licence)
-
Renmin University of China at TRECVID 2022: Improving Video Search by Feature Fusion and Negation Understanding 28 Nov 2022 · 0 repositories · arXiv:2211.15039