Methods › Computer Vision › Vision and Language Pre-Trained Models › CLIP › Papers, page 22
Contrastive Language-Image Pre-training
CLIP
Papers archive 2025-07-28
archive papers tagged: 3,094 · with a code link: 1,617 · where Syntology ran a sample: 649 (554 with a run with no instrument failure, 95 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (649 of 3,094 tagged: 554 with a run with no instrument failure, 95 where every run was a failure of Syntology's instrument)
Page 22 of 31: papers 2,101 to 2,200 of 3,094, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Adversarial Attacks on Foundational Vision Models 28 Aug 2023 · 0 repositories · arXiv:2308.14597
-
Do the Frankenstein, or how to achieve better out-of-distribution performance with manifold mixing model soup 28 Aug 2023 · 0 repositories · arXiv:2309.08610
-
Exploring the Transfer Learning Capabilities of CLIP in Domain Generalization for Diabetic Retinopathy 27 Aug 2023 · 1 repository · arXiv:2308.14212
-
Prompting Visual-Language Models for Dynamic Facial Expression Recognition 25 Aug 2023 · 1 repository · arXiv:2308.13382
-
Parameter-Efficient Transfer Learning for Remote Sensing Image-Text Retrieval 24 Aug 2023 · 1 repository · arXiv:2308.12509Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
PartSeg: Few-shot Part Segmentation via Part-aware Prompt Learning 24 Aug 2023 · 0 repositories · arXiv:2308.12757
-
PromptMRG: Diagnosis-Driven Prompts for Medical Report Generation 24 Aug 2023 · 1 repository · arXiv:2308.12604Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample)
-
Scenimefy: Learning to Craft Anime Scene via Semi-Supervised Image-to-Image Translation 24 Aug 2023 · 1 repository · arXiv:2308.12968
-
Realistic Unsupervised CLIP Fine-tuning with Universal Entropy Optimization 24 Aug 2023 · 1 repository · arXiv:2308.12919
-
Towards Realistic Zero-Shot Classification via Self Structural Semantic Alignment 24 Aug 2023 · 1 repository · arXiv:2308.12960
-
Blending-NeRF: Text-Driven Localized Editing in Neural Radiance Fields 23 Aug 2023 · 0 repositories · arXiv:2308.11974
-
CLIPN for Zero-Shot OOD Detection: Teaching CLIP to Say No 23 Aug 2023 · 1 repository · arXiv:2308.12213Syntology official (archive's flag): 10 ran · 10 ran (of which 1 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 2 where Syntology's instrument failed) · 6 unverified (of 16 harvested samples) · 2 pointer-only (licence)
-
Classification of the lunar surface pattern by AI architectures: Does AI see a rabbit in the Moon? 22 Aug 2023 · 0 repositories · arXiv:2308.11107
-
CLIP Multi-modal Hashing: A new baseline CLIPMH 22 Aug 2023 · 0 repositories · arXiv:2308.11797
-
Composed Image Retrieval using Contrastive Learning and Task-oriented CLIP-based Features 22 Aug 2023 · 2 repositories · arXiv:2308.11485Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples)
-
GOPro: Generate and Optimize Prompts in CLIP using Self-Supervised Learning 22 Aug 2023 · 1 repository · arXiv:2308.11605
-
Knowledge-Aware Prompt Tuning for Generalizable Vision-Language Models 22 Aug 2023 · 0 repositories · arXiv:2308.11186
-
LCCo: Lending CLIP to Co-Segmentation 22 Aug 2023 · 0 repositories · arXiv:2308.11506
-
Opening the Vocabulary of Egocentric Actions 22 Aug 2023 · 1 repository · arXiv:2308.11488
-
Random Word Data Augmentation with CLIP for Zero-Shot Anomaly Detection 22 Aug 2023 · 0 repositories · arXiv:2308.11119
-
Unsupervised Prototype Adapter for Vision-Language Models 22 Aug 2023 · 0 repositories · arXiv:2308.11507
-
VadCLIP: Adapting Vision-Language Models for Weakly Supervised Video Anomaly Detection 22 Aug 2023 · 1 repository · arXiv:2308.11681Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples)
-
ViLLA: Fine-Grained Vision-Language Representation Learning from Real-World Data 22 Aug 2023 · 1 repository · arXiv:2308.11194Syntology official (archive's flag): 1 ran · 1 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified; the one sample that ran constructed an object rather than computing a result (of 4 harvested samples)
-
An Examination of the Compositionality of Large Generative Vision-Language Models 21 Aug 2023 · 1 repository · arXiv:2308.10509
-
Improving Diversity in Zero-Shot GAN Adaptation with Semantic Variations 21 Aug 2023 · 0 repositories · arXiv:2308.10554
-
SCULPT: Shape-Conditioned Unpaired Learning of Pose-dependent Clothed and Textured Human Meshes 21 Aug 2023 · 0 repositories · arXiv:2308.10638
-
Turning a CLIP Model into a Scene Text Spotter 21 Aug 2023 · 1 repository · arXiv:2308.10408
-
UnLoc: A Unified Framework for Video Localization Tasks 21 Aug 2023 · 1 repository · arXiv:2308.11062
-
Generic Attention-model Explainability by Weighted Relevance Accumulation 20 Aug 2023 · 0 repositories · arXiv:2308.10240
-
An Empirical Study of CLIP for Text-based Person Search 19 Aug 2023 · 1 repository · arXiv:2308.10045
-
Finding emergence in data by maximizing effective information 19 Aug 2023 · 0 repositories · arXiv:2308.09952
-
Audio-Visual Glance Network for Efficient Video Recognition 18 Aug 2023 · 0 repositories · arXiv:2308.09322
-
DiffDis: Empowering Generative Diffusion Model with Cross-Modal Discrimination Capability 18 Aug 2023 · 0 repositories · arXiv:2308.09306
-
Label-Free Event-based Object Recognition via Joint Learning with Image Reconstruction from Events 18 Aug 2023 · 1 repository · arXiv:2308.09383Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 1 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples) · 1 pointer-only (licence)
-
Point Contrastive Prediction with Semantic Clustering for Self-Supervised Learning on Point Cloud Videos 18 Aug 2023 · 0 repositories · arXiv:2308.09247
-
V2A-Mapper: A Lightweight Solution for Vision-to-Audio Generation by Connecting Foundation Models 18 Aug 2023 · 1 repository · arXiv:2308.09300
-
EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding 17 Aug 2023 · 1 repository · arXiv:2308.09126Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Text-Only Training for Visual Storytelling 17 Aug 2023 · 0 repositories · arXiv:2308.08881
-
Visually-Aware Context Modeling for News Image Captioning 16 Aug 2023 · 1 repository · arXiv:2308.08325
-
Exploring Transfer Learning in Medical Image Segmentation using Vision-Language Models 15 Aug 2023 · 1 repository · arXiv:2308.07706
-
Prompt Switch: Efficient CLIP Adaptation for Text-Video Retrieval 15 Aug 2023 · 1 repository · arXiv:2308.07648Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 1 where Syntology's instrument failed) · 12 unverified (of 16 harvested samples) · 16 pointer-only (licence)
-
StyleDiffusion: Controllable Disentangled Style Transfer via Diffusion Models 15 Aug 2023 · 0 repositories · arXiv:2308.07863
-
Visual and Textual Prior Guided Mask Assemble for Few-Shot Segmentation and Beyond 15 Aug 2023 · 0 repositories · arXiv:2308.07539
-
AdvCLIP: Downstream-agnostic Adversarial Examples in Multimodal Contrastive Learning 14 Aug 2023 · 1 repository · arXiv:2308.07026Syntology official (archive's flag): 15 ran · 15 ran (of which 0 constructed an object rather than computing a result; 12 with no instrument failure: 1 honoured, 1 violated, 10 with no contract checked; 3 where Syntology's instrument failed) · 1 unverified (of 16 harvested samples) · 5 pointer-only (licence)
-
ICPC: Instance-Conditioned Prompting with Contrastive Learning for Semantic Segmentation 14 Aug 2023 · 0 repositories · arXiv:2308.07078
-
Neural Categorical Priors for Physics-Based Character Control 14 Aug 2023 · 0 repositories · arXiv:2308.07200
-
Semantify: Simplifying the Control of 3D Morphable Models using CLIP 14 Aug 2023 · 1 repository · arXiv:2308.07415
-
UniBrain: Unify Image Reconstruction and Captioning All in One Diffusion Model from Human Brain Activity 14 Aug 2023 · 0 repositories · arXiv:2308.07428
-
Diverse Data Augmentation with Diffusions for Effective Test-time Prompt Tuning 11 Aug 2023 · 1 repository · arXiv:2308.06038
-
Multimodality and Attention Increase Alignment in Natural Language Prediction Between Humans and Computational Models 11 Aug 2023 · 0 repositories · arXiv:2308.06035
-
AD-CLIP: Adapting Domains in Prompt Space Using CLIP 10 Aug 2023 · 1 repository · arXiv:2308.05659Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 2 unverified (of 5 harvested samples) · 2 pointer-only (licence)
-
TCSloT: Text Guided 3D Context and Slope Aware Triple Network for Dental Implant Position Prediction 10 Aug 2023 · 0 repositories · arXiv:2308.05355
-
Seeing in Flowing: Adapting CLIP for Action Recognition with Motion Prompts Learning 9 Aug 2023 · 0 repositories · arXiv:2308.04828
-
MindDiffuser: Controlled Image Reconstruction from Human Brain Activity with Semantic and Structural Diffusion 8 Aug 2023 · 1 repository · arXiv:2308.04249Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples)
-
The Five-Dollar Model: Generating Game Maps and Sprites from Sentence Embeddings 8 Aug 2023 · 1 repository · arXiv:2308.04052
-
Distributionally Robust Classification on a Data Budget 7 Aug 2023 · 1 repository · arXiv:2308.03821Syntology official: harvested, nothing ran · 0 ran · 2 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
EventBind: Learning a Unified Representation to Bind Them All for Event-based Open-world Understanding 6 Aug 2023 · 0 repositories · arXiv:2308.03135
-
Photorealistic and Identity-Preserving Image-Based Emotion Manipulation with Latent Diffusion Models 6 Aug 2023 · 1 repository · arXiv:2308.03183
-
Improving Generalization of Image Captioning with Unsupervised Prompt Learning 5 Aug 2023 · 0 repositories · arXiv:2308.02862
-
Convolutions Die Hard: Open-Vocabulary Segmentation with Single Frozen Convolutional CLIP 4 Aug 2023 · 1 repository · arXiv:2308.02487Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
ReCLIP: Refine Contrastive Language Image Pre-Training with Source Free Domain Adaptation 4 Aug 2023 · 1 repository · arXiv:2308.03793Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample)
-
MA-FSAR: Multimodal Adaptation of CLIP for Few-Shot Action Recognition 3 Aug 2023 · 0 repositories · arXiv:2308.01532
-
PerceptionCLIP: Visual Classification by Inferring and Conditioning on Contexts 2 Aug 2023 · 1 repository · arXiv:2308.01313Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 4 where Syntology's instrument failed) · 8 unverified (of 12 harvested samples) · 5 pointer-only (licence)
-
Rethinking Similarity Search: Embracing Smarter Mechanisms over Smarter Data 2 Aug 2023 · 0 repositories · arXiv:2308.00909
-
Detecting Cloud Presence in Satellite Images Using the RGB-based CLIP Vision-Language Model 1 Aug 2023 · 0 repositories · arXiv:2308.00541
-
CDUL: CLIP-Driven Unsupervised Learning for Multi-Label Image Classification 31 Jul 2023 · 1 repository · arXiv:2307.16634
-
Guiding Image Captioning Models Toward More Specific Captions 31 Jul 2023 · 0 repositories · arXiv:2307.16686
-
Open-Set Domain Adaptation with Visual-Language Foundation Models 30 Jul 2023 · 0 repositories · arXiv:2307.16204
-
Sat2Cap: Mapping Fine-Grained Textual Descriptions from Satellite Images 29 Jul 2023 · 0 repositories · arXiv:2307.15904
-
CLIP Brings Better Features to Visual Aesthetics Learners 28 Jul 2023 · 0 repositories · arXiv:2307.15640
-
Cross-Modal Concept Learning and Inference for Vision-Language Models 28 Jul 2023 · 0 repositories · arXiv:2307.15460
-
Distilled Feature Fields Enable Few-Shot Language-Guided Manipulation 27 Jul 2023 · 1 repository · arXiv:2308.07931Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples)
-
Learning Multi-modal Representations by Watching Hundreds of Surgical Video Lectures 27 Jul 2023 · 2 repositories · arXiv:2307.15220Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
PromptStyler: Prompt-driven Style Generation for Source-free Domain Generalization 27 Jul 2023 · 1 repository · arXiv:2307.15199Syntology 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Regularized Mask Tuning: Uncovering Hidden Knowledge in Pre-trained Vision-Language Models 27 Jul 2023 · 0 repositories · arXiv:2307.15049
-
Self-Supervised Visual Acoustic Matching 27 Jul 2023 · 0 repositories · arXiv:2307.15064
-
ECO: Ensembling Context Optimization for Vision-Language Models 26 Jul 2023 · 0 repositories · arXiv:2307.14063
-
Jailbreak in pieces: Compositional Adversarial Attacks on Multi-Modal Language Models 26 Jul 2023 · 0 repositories · arXiv:2307.14539
-
Benchmarking and Analyzing Generative Data for Visual Recognition 25 Jul 2023 · 0 repositories · arXiv:2307.13697
-
The Visual Language of Fabrics 25 Jul 2023 · 0 repositories · arXiv:2307.13681
-
CLIP-KD: An Empirical Study of CLIP Model Distillation 24 Jul 2023 · 1 repository · arXiv:2307.12732Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 1 violated, 4 with no contract checked; 4 where Syntology's instrument failed) · 9 unverified (of 18 harvested samples) · 18 pointer-only (licence)
-
Does Progress On Object Recognition Benchmarks Improve Real-World Generalization? 24 Jul 2023 · 0 repositories · arXiv:2307.13136
-
Interpolating between Images with Diffusion Models 24 Jul 2023 · 1 repository · arXiv:2307.12560Syntology 5 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 6 harvested samples) · 2 pointer-only (licence)
-
Geometry-Aware Adaptation for Pretrained Models 23 Jul 2023 · 0 repositories · arXiv:2307.12226
-
SCRAPS: Speech Contrastive Representations of Acoustic and Phonetic Spaces 23 Jul 2023 · 0 repositories · arXiv:2307.12445
-
Why Is Prompt Tuning for Vision-Language Models Robust to Noisy Labels? 22 Jul 2023 · 1 repository · arXiv:2307.11978Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 2 unverified (of 5 harvested samples) · 4 pointer-only (licence)
-
Batch Clipping and Adaptive Layerwise Clipping for Differential Private Stochastic Gradient Descent 21 Jul 2023 · 0 repositories · arXiv:2307.11939
-
Enhancing CLIP with GPT-4: Harnessing Visual Descriptions as Prompts 21 Jul 2023 · 1 repository · arXiv:2307.11661Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 4 harvested samples) · 1 pointer-only (licence)
-
FaceCLIPNeRF: Text-driven 3D Face Manipulation using Deformable Neural Radiance Fields 21 Jul 2023 · 0 repositories · arXiv:2307.11418
-
GIST: Generating Image-Specific Text for Fine-grained Object Classification 21 Jul 2023 · 1 repository · arXiv:2307.11315
-
Tuning Pre-trained Model via Moment Probing 21 Jul 2023 · 1 repository · arXiv:2307.11342Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 1 honoured, 0 violated, 2 with no contract checked; 5 where Syntology's instrument failed) · 5 unverified (of 13 harvested samples) · 13 pointer-only (licence)
-
Identifying Interpretable Subspaces in Image Representations 20 Jul 2023 · 2 repositories · arXiv:2307.10504Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Language-based Action Concept Spaces Improve Video Self-Supervised Learning 20 Jul 2023 · 0 repositories · arXiv:2307.10922
-
Reference-based Painterly Inpainting via Diffusion: Crossing the Wild Reference Domain Gap 20 Jul 2023 · 0 repositories · arXiv:2307.10584
-
UP-DP: Unsupervised Prompt Learning for Data Pre-Selection with Vision-Language Models 20 Jul 2023 · 0 repositories · arXiv:2307.11227
-
Findings of Factify 2: Multimodal Fake News Detection 19 Jul 2023 · 0 repositories · arXiv:2307.10475
-
Improving Multimodal Datasets with Image Captioning 19 Jul 2023 · 0 repositories · arXiv:2307.10350
-
LDP: Language-driven Dual-Pixel Image Defocus Deblurring Network 19 Jul 2023 · 0 repositories · arXiv:2307.09815
-
Watch out Venomous Snake Species: A Solution to SnakeCLEF2023 19 Jul 2023 · 1 repository · arXiv:2307.09748
-
Distilling Knowledge from Text-to-Image Generative Models Improves Visio-Linguistic Reasoning in CLIP 18 Jul 2023 · 0 repositories · arXiv:2307.09233