Methods › Computer Vision › Vision and Language Pre-Trained Models › CLIP › Papers, page 26
Contrastive Language-Image Pre-training
CLIP
Papers archive 2025-07-28
archive papers tagged: 3,094 · with a code link: 1,617 · where Syntology ran a sample: 649 (554 with a run with no instrument failure, 95 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (649 of 3,094 tagged: 554 with a run with no instrument failure, 95 where every run was a failure of Syntology's instrument)
Page 26 of 31: papers 2,501 to 2,600 of 3,094, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Automatic Geo-alignment of Artwork in Children's Story Books 16 Mar 2023 · 0 repositories · arXiv:2304.01204
-
Global Knowledge Calibration for Fast Open-Vocabulary Segmentation 16 Mar 2023 · 1 repository · arXiv:2303.09181
-
GridCLIP: One-Stage Object Detection by Grid-Level CLIP Representation Learning 16 Mar 2023 · 0 repositories · arXiv:2303.09252
-
LERF: Language Embedded Radiance Fields 16 Mar 2023 · 5 repositories · arXiv:2303.09553
-
MultiModal Bias: Introducing a Framework for Stereotypical Bias Assessment beyond Gender and Race in Vision Language Models 16 Mar 2023 · 1 repository · arXiv:2303.12734
-
SpectralCLIP: Preventing Artifacts in Text-Guided Style Transfer from a Spectral Perspective 16 Mar 2023 · 1 repository · arXiv:2303.09270
-
StylerDALLE: Language-Guided Style Transfer Using a Vector-Quantized Tokenizer of a Large-Scale Generative Model 16 Mar 2023 · 1 repository · arXiv:2303.09268Syntology official (archive's flag): 9 ran · 9 ran (of which 9 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified; every one of the 9 samples that ran constructed an object rather than computing a result (of 10 harvested samples) · 10 pointer-only (licence)
-
TemporalMaxer: Maximize Temporal Context with only Max Pooling for Temporal Action Localization 16 Mar 2023 · 1 repository · arXiv:2303.09055
-
VideoFusion: Decomposed Diffusion Models for High-Quality Video Generation 15 Mar 2023 · 2 repositories · arXiv:2303.08320
-
DeepMIM: Deep Supervision for Masked Image Modeling 15 Mar 2023 · 1 repository · arXiv:2303.08817
-
Highly Personalized Text Embedding for Image Manipulation by Stable Diffusion 15 Mar 2023 · 0 repositories · arXiv:2303.08767
-
PR-MCS: Perturbation Robust Metric for MultiLingual Image Captioning 15 Mar 2023 · 0 repositories · arXiv:2303.08389
-
Generation-Guided Multi-Level Unified Network for Video Grounding 14 Mar 2023 · 0 repositories · arXiv:2303.07748
-
Robust Contrastive Language-Image Pre-training against Data Poisoning and Backdoor Attacks 13 Mar 2023 · 1 repository · arXiv:2303.06854Syntology official: no sample here; runs from other or unrecorded repositories · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Accommodating Audio Modality in CLIP for Multimodal Processing 12 Mar 2023 · 1 repository · arXiv:2303.06591
-
One Transformer Fits All Distributions in Multi-Modal Diffusion at Scale 12 Mar 2023 · 3 repositories · arXiv:2303.06555Syntology official (archive's flag): 1 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Preventing Zero-Shot Transfer Degradation in Continual Learning of Vision-Language Models 12 Mar 2023 · 2 repositories · arXiv:2303.06628
-
Towards Universal Vision-language Omni-supervised Segmentation 12 Mar 2023 · 0 repositories · arXiv:2303.06547
-
DeltaEdit: Exploring Text-free Training for Text-Driven Image Manipulation 11 Mar 2023 · 1 repository · arXiv:2303.06285
-
Enabling Calibration In The Zero-Shot Inference of Large Vision-Language Models 11 Mar 2023 · 0 repositories · arXiv:2303.12748
-
Adapting Contrastive Language-Image Pretrained (CLIP) Models for Out-of-Distribution Detection 10 Mar 2023 · 1 repository · arXiv:2303.05828
-
Iterative Few-shot Semantic Segmentation from Image Label Text 10 Mar 2023 · 1 repository · arXiv:2303.05646
-
Object-Aware Distillation Pyramid for Open-Vocabulary Object Detection 10 Mar 2023 · 1 repository · arXiv:2303.05892
-
Open-Ended Medical Visual Question Answering Through Prefix Tuning of Language Models 10 Mar 2023 · 1 repository · arXiv:2303.05977Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Detecting Images Generated by Diffusers 9 Mar 2023 · 1 repository · arXiv:2303.05275
-
Mimic before Reconstruct: Enhancing Masked Autoencoders with Feature Mimicking 9 Mar 2023 · 1 repository · arXiv:2303.05475
-
Multimodal Parameter-Efficient Few-Shot Class Incremental Learning 8 Mar 2023 · 1 repository · arXiv:2303.04751
-
Open-Vocabulary Panoptic Segmentation with Text-to-Image Diffusion Models 8 Mar 2023 · 1 repository · arXiv:2303.04803
-
A Comprehensive Survey of AI-Generated Content (AIGC): A History of Generative AI from GAN to ChatGPT 7 Mar 2023 · 1 repository · arXiv:2303.04226
-
CLIP-Layout: Style-Consistent Indoor Scene Synthesis with Semantic Furniture Embedding 7 Mar 2023 · 1 repository · arXiv:2303.03565
-
MOSO: Decomposing MOtion, Scene and Object for Video Prediction 7 Mar 2023 · 2 repositories · arXiv:2303.03684
-
Can an Embodied Agent Find Your "Cat-shaped Mug"? LLM-Guided Exploration for Zero-Shot Object Navigation 6 Mar 2023 · 1 repository · arXiv:2303.03480
-
CleanCLIP: Mitigating Data Poisoning Attacks in Multimodal Contrastive Learning 6 Mar 2023 · 1 repository · arXiv:2303.03323Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 3 harvested samples)
-
CLIP-guided Prototype Modulating for Few-shot Action Recognition 6 Mar 2023 · 1 repository · arXiv:2303.02982Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 4 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples) · 3 pointer-only (licence)
-
DeCap: Decoding CLIP Latents for Zero-Shot Captioning via Text-Only Training 6 Mar 2023 · 1 repository · arXiv:2303.03032
-
HiCLIP: Contrastive Language-Image Pretraining with Hierarchy-aware Attention 6 Mar 2023 · 1 repository · arXiv:2303.02995
-
IPA-CLIP: Integrating Phonetic Priors into Vision and Language Pretraining 6 Mar 2023 · 0 repositories · arXiv:2303.03144
-
Video Question Answering Using CLIP-Guided Visual-Text Attention 6 Mar 2023 · 0 repositories · arXiv:2303.03131
-
Improving Audio-Visual Video Parsing with Pseudo Visual Labels 4 Mar 2023 · 0 repositories · arXiv:2303.02344
-
Prompt, Generate, then Cache: Cascade of Foundation Models makes Strong Few-shot Learners 3 Mar 2023 · 3 repositories · arXiv:2303.02151Syntology official: not harvested · 0 ran · 1 unverified (of 1 harvested sample)
-
INO at Factify 2: Structure Coherence based Multi-Modal Fact Verification 2 Mar 2023 · 1 repository · arXiv:2303.01510
-
BiomedCLIP: a multimodal biomedical foundation model pretrained from fifteen million scientific image-text pairs 2 Mar 2023 · 5 repositories · arXiv:2303.00915
-
Zero-Shot Text-to-Parameter Translation for Game Character Auto-Creation 2 Mar 2023 · 0 repositories · arXiv:2303.01311
-
CLIPER: A Unified Vision-Language Framework for In-the-Wild Facial Expression Recognition 1 Mar 2023 · 0 repositories · arXiv:2303.00193
-
An Effective Crop-Paste Pipeline for Few-shot Object Detection 28 Feb 2023 · 0 repositories · arXiv:2302.14452
-
TextIR: A Simple Framework for Text-based Editable Image Restoration 28 Feb 2023 · 0 repositories · arXiv:2302.14736
-
Turning a CLIP Model into a Scene Text Detector 28 Feb 2023 · 1 repository · arXiv:2302.14338
-
A Language-Guided Benchmark for Weakly Supervised Open Vocabulary Semantic Segmentation 27 Feb 2023 · 1 repository · arXiv:2302.14163
-
FedCLIP: Fast Generalization and Personalization for CLIP in Federated Learning 27 Feb 2023 · 1 repository · arXiv:2302.13485Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 12 harvested samples)
-
Internet Explorer: Targeted Representation Learning on the Open Web 27 Feb 2023 · 1 repository · arXiv:2302.14051
-
Vid2Seq: Large-Scale Pretraining of a Visual Language Model for Dense Video Captioning 27 Feb 2023 · 3 repositories · arXiv:2302.14115
-
Agile Modeling: From Concept to Classifier in Minutes 25 Feb 2023 · 0 repositories · arXiv:2302.12948
-
BrainCLIP: Bridging Brain and Visual-Linguistic Representation Via CLIP for Generic Natural Visual Stimulus Decoding 25 Feb 2023 · 1 repository · arXiv:2302.12971
-
A framework for benchmarking class-out-of-distribution detection and its application to ImageNet 23 Feb 2023 · 1 repository · arXiv:2302.11893Syntology official (archive's flag): 13 ran · 13 ran (of which 0 constructed an object rather than computing a result; 13 with no instrument failure: 0 honoured, 0 violated, 13 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 15 harvested samples)
-
Controlled and Conditional Text to Image Generation with Diffusion Prior 23 Feb 2023 · 0 repositories · arXiv:2302.11710
-
Side Adapter Network for Open-Vocabulary Semantic Segmentation 23 Feb 2023 · 3 repositories · arXiv:2302.12242Syntology official (archive's flag): 1 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 9 harvested samples)
-
Teaching CLIP to Count to Ten 23 Feb 2023 · 1 repository · arXiv:2302.12066Syntology 0 ran · 1 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Test-Time Distribution Normalization for Contrastively Learned Vision-language Models 22 Feb 2023 · 2 repositories · arXiv:2302.11084Syntology official: no sample here; runs from other or unrecorded repositories · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Open-domain Visual Entity Recognition: Towards Recognizing Millions of Wikipedia Entities 22 Feb 2023 · 2 repositories · arXiv:2302.11154Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 3 unverified (of 4 harvested samples)
-
X-TRA: Improving Chest X-ray Tasks with Cross-Modal Retrieval Augmentation 22 Feb 2023 · 0 repositories · arXiv:2302.11352
-
A Picture May Be Worth a Thousand Lives: An Interpretable Artificial Intelligence Strategy for Predictions of Suicide Risk from Social Media Images 19 Feb 2023 · 0 repositories · arXiv:2302.09488
-
StyLIP: Multi-Scale Style-Conditioned Prompt Learning for CLIP-based Domain Generalization 18 Feb 2023 · 0 repositories · arXiv:2302.09251
-
Towards Efficient Visual Adaption via Structural Re-parameterization 16 Feb 2023 · 1 repository · arXiv:2302.08106Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
PRedItOR: Text Guided Image Editing with Diffusion Prior 15 Feb 2023 · 0 repositories · arXiv:2302.07979
-
A Modern Look at the Relationship between Sharpness and Generalization 14 Feb 2023 · 1 repository · arXiv:2302.07011
-
Actional Atomic-Concept Learning for Demystifying Vision-Language Navigation 13 Feb 2023 · 0 repositories · arXiv:2302.06072
-
VITR: Augmenting Vision Transformers with Relation-Focused Learning for Cross-Modal Information Retrieval 13 Feb 2023 · 0 repositories · arXiv:2302.06350
-
NYCU-TWO at Memotion 3: Good Foundation, Good Teacher, then you have Good Meme Analysis 13 Feb 2023 · 0 repositories · arXiv:2302.06078
-
Paparazzi: A Deep Dive into the Capabilities of Language and Vision Models for Grounding Viewpoint Descriptions 13 Feb 2023 · 0 repositories · arXiv:2302.10282
-
Understanding Multimodal Contrastive Learning and Incorporating Unpaired Data 13 Feb 2023 · 1 repository · arXiv:2302.06232
-
Differentiable Outlier Detection Enable Robust Deep Multimodal Analysis 11 Feb 2023 · 1 repository · arXiv:2302.05608
-
Auditing Gender Presentation Differences in Text-to-Image Models 7 Feb 2023 · 1 repository · arXiv:2302.03675Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Diversity is Definitely Needed: Improving Model-Agnostic Zero-shot Classification via Stable Diffusion 7 Feb 2023 · 1 repository · arXiv:2302.03298Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
CHiLS: Zero-Shot Image Classification with Hierarchical Label Sets 6 Feb 2023 · 1 repository · arXiv:2302.02551Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 11 harvested samples) · 1 pointer-only (licence)
-
LexLIP: Lexicon-Bottlenecked Language-Image Pre-Training for Large-Scale Image-Text Retrieval 6 Feb 2023 · 1 repository · arXiv:2302.02908
-
MOSE: A New Dataset for Video Object Segmentation in Complex Scenes 3 Feb 2023 · 1 repository · arXiv:2302.01872Syntology official (archive's flag): 6 ran · 9 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 2 honoured, 1 violated, 3 with no contract checked; 3 where Syntology's instrument failed) · 3 unverified (of 12 harvested samples) · 6 pointer-only (licence)
-
CLIPood: Generalizing CLIP to Out-of-Distributions 2 Feb 2023 · 1 repository · arXiv:2302.00864Syntology official (archive's flag): 1 ran · 1 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified; the one sample that ran constructed an object rather than computing a result (of 1 harvested sample)
-
Learning Generalized Zero-Shot Learners for Open-Domain Image Geolocalization 1 Feb 2023 · 1 repository · arXiv:2302.00275
-
Open-VCLIP: Transforming CLIP to an Open-vocabulary Video Model via Interpolated Weight Optimization 1 Feb 2023 · 1 repository · arXiv:2302.00624Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Automated Time-frequency Domain Audio Crossfades using Graph Cuts 31 Jan 2023 · 0 repositories · arXiv:2301.13380
-
Zero3D: Semantic-Driven Multi-Category 3D Shape Generation 31 Jan 2023 · 0 repositories · arXiv:2301.13591
-
A Bias-Accuracy-Privacy Trilemma for Statistical Estimation 30 Jan 2023 · 0 repositories · arXiv:2301.13334
-
GALIP: Generative Adversarial CLIPs for Text-to-Image Synthesis 30 Jan 2023 · 2 repositories · arXiv:2301.12959Syntology official (archive's flag): 6 ran · 7 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 1 where Syntology's instrument failed) · 4 unverified (of 11 harvested samples) · 2 pointer-only (licence)
-
STAIR: Learning Sparse Text and Image Representation in Grounded Tokens 30 Jan 2023 · 0 repositories · arXiv:2301.13081
-
ZegOT: Zero-shot Segmentation Through Optimal Transport of Text Prompts 28 Jan 2023 · 1 repository · arXiv:2301.12171
-
Discovering and Mitigating Visual Biases through Keyword Explanation 26 Jan 2023 · 1 repository · arXiv:2301.11104
-
Joint action loss for proximal policy optimization 26 Jan 2023 · 1 repository · arXiv:2301.10919
-
Revisiting Temporal Modeling for CLIP-based Image-to-Video Knowledge Transferring 26 Jan 2023 · 1 repository · arXiv:2301.11116
-
Vision-Language Models Performing Zero-Shot Tasks Exhibit Gender-based Disparities 26 Jan 2023 · 0 repositories · arXiv:2301.11100
-
Towards Arbitrary Text-driven Image Manipulation via Space Alignment 25 Jan 2023 · 0 repositories · arXiv:2301.10670
-
OvarNet: Towards Open-vocabulary Object Attribute Recognition 23 Jan 2023 · 1 repository · arXiv:2301.09506
-
Exploring the Synergy Between Vision-Language Pretraining and ChatGPT for Artwork Captioning: A Preliminary Study 21 Jan 2023 · 1 repository
-
MV-Adapter: Multimodal Video Transfer Learning for Video Text Retrieval 19 Jan 2023 · 1 repository · arXiv:2301.07868
-
Masked Autoencoding Does Not Help Natural Language Supervision at Scale 19 Jan 2023 · 0 repositories · arXiv:2301.07836
-
CLIPTER: Looking at the Bigger Picture in Scene Text Recognition 18 Jan 2023 · 0 repositories · arXiv:2301.07464
-
Face Recognition in the age of CLIP & Billion image datasets 18 Jan 2023 · 0 repositories · arXiv:2301.07315
-
Joint Representation Learning for Text and 3D Point Cloud 18 Jan 2023 · 0 repositories · arXiv:2301.07584
-
Learning Customized Visual Models with Retrieval-Augmented Knowledge 17 Jan 2023 · 1 repository · arXiv:2301.07094Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 1 violated, 4 with no contract checked; 4 where Syntology's instrument failed) · 8 unverified (of 17 harvested samples) · 8 pointer-only (licence)
-
RILS: Masked Visual Reconstruction in Language Semantic Space 17 Jan 2023 · 1 repository · arXiv:2301.06958Syntology official (archive's flag): 1 ran · 1 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified; the one sample that ran constructed an object rather than computing a result (of 2 harvested samples) · 2 pointer-only (licence)
-
USER: Unified Semantic Enhancement with Momentum Contrast for Image-Text Retrieval 17 Jan 2023 · 1 repository · arXiv:2301.06844