Methods › Computer Vision › Vision and Language Pre-Trained Models › CLIP › Papers, page 14
Contrastive Language-Image Pre-training
CLIP
Papers archive 2025-07-28
archive papers tagged: 3,094 · with a code link: 1,617 · where Syntology ran a sample: 649 (554 with a run with no instrument failure, 95 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (649 of 3,094 tagged: 554 with a run with no instrument failure, 95 where every run was a failure of Syntology's instrument)
Page 14 of 31: papers 1,301 to 1,400 of 3,094, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Finding Shared Decodable Concepts and their Negations in the Brain 27 May 2024 · 0 repositories · arXiv:2405.17663
-
CapS-Adapter: Caption-based MultiModal Adapter in Zero-Shot Classification 26 May 2024 · 1 repository · arXiv:2405.16591
-
Disentangling Foreground and Background Motion for Enhanced Realism in Human Video Generation 26 May 2024 · 0 repositories · arXiv:2405.16393
-
Accelerating Transformers with Spectrum-Preserving Token Merging 25 May 2024 · 1 repository · arXiv:2405.16148Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 2 honoured, 0 violated, 0 with no contract checked; 4 where Syntology's instrument failed) · 1 unverified (of 7 harvested samples) · 7 pointer-only (licence)
-
An Empirical Study of Excitation and Aggregation Design Adaptions in CLIP4Clip for Video-Text Retrieval 25 May 2024 · 0 repositories · arXiv:2406.01604
-
Dual-Adapter: Training-free Dual Adaptation for Few-shot Out-of-Distribution Detection 25 May 2024 · 0 repositories · arXiv:2405.16146
-
Enhancing Near OOD Detection in Prompt Learning: Maximum Gains, Minimal Costs 25 May 2024 · 0 repositories · arXiv:2405.16091
-
How Well Do Deep Learning Models Capture Human Concepts? The Case of the Typicality Effect 25 May 2024 · 0 repositories · arXiv:2405.16128
-
Streaming Long Video Understanding with Large Language Models 25 May 2024 · 0 repositories · arXiv:2405.16009
-
Underwater Image Enhancement by Diffusion Model with Customized CLIP-Classifier 25 May 2024 · 1 repository · arXiv:2405.16214
-
BDetCLIP: Multimodal Prompting Contrastive Test-Time Backdoor Detection 24 May 2024 · 0 repositories · arXiv:2405.15269
-
CLIP model is an Efficient Online Lifelong Learner 24 May 2024 · 1 repository · arXiv:2405.15155
-
Learning Invariant Causal Mechanism from Vision-Language Models 24 May 2024 · 0 repositories · arXiv:2405.15289Syntology 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
ArtWeaver: Advanced Dynamic Style Integration via Diffusion Model 24 May 2024 · 0 repositories · arXiv:2405.15287
-
A Lost Opportunity for Vision-Language Models: A Comparative Study of Online Test-Time Adaptation for Vision-Language Models 23 May 2024 · 1 repository · arXiv:2405.14977
-
CLIPScope: Enhancing Zero-Shot OOD Detection with Bayesian Scoring 23 May 2024 · 1 repository · arXiv:2405.14737Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Concept Visualization: Explaining the CLIP Multi-modal Embedding Using WordNet 23 May 2024 · 1 repository · arXiv:2405.14563
-
Designing A Sustainable Marine Debris Clean-up Framework without Human Labels 23 May 2024 · 1 repository · arXiv:2405.14815
-
Efficiency for Free: Ideal Data Are Transportable Representations 23 May 2024 · 1 repository · arXiv:2405.14669
-
Harmony: A Joint Self-Supervised and Weakly-Supervised Framework for Learning General Purpose Visual Representations 23 May 2024 · 1 repository · arXiv:2405.14239
-
TUNI: A Textual Unimodal Detector for Identity Inference in CLIP Models 23 May 2024 · 0 repositories · arXiv:2405.14517
-
Learning Multi-dimensional Human Preference for Text-to-Image Generation 23 May 2024 · 1 repository · arXiv:2405.14705
-
Leveraging Semantic Segmentation Masks with Embeddings for Fine-Grained Form Classification 23 May 2024 · 0 repositories · arXiv:2405.14162
-
Pre-Trained Vision-Language Models as Partial Annotators 23 May 2024 · 0 repositories · arXiv:2406.18550
-
Text-to-Model: Text-Conditioned Neural Network Diffusion for Train-Once-for-All Personalization 23 May 2024 · 0 repositories · arXiv:2405.14132
-
Towards Cross-modal Backward-compatible Representation Learning for Vision-Language Models 23 May 2024 · 0 repositories · arXiv:2405.14715
-
Tuning-free Universally-Supervised Semantic Segmentation 23 May 2024 · 0 repositories · arXiv:2405.14294
-
GMMFormer v2: An Uncertainty-aware Framework for Partially Relevant Video Retrieval 22 May 2024 · 1 repository · arXiv:2405.13824Syntology official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 10 harvested samples) · 10 pointer-only (licence)
-
Gradient Projection For Continual Parameter-Efficient Tuning 22 May 2024 · 0 repositories · arXiv:2405.13383
-
Monocular Gaussian SLAM with Language Extended Loop Closure 22 May 2024 · 0 repositories · arXiv:2405.13748
-
Refining Skewed Perceptions in Vision-Language Models through Visual Representations 22 May 2024 · 0 repositories · arXiv:2405.14030
-
TOPA: Extending Large Language Models for Video Understanding via Text-Only Pre-Alignment 22 May 2024 · 1 repository · arXiv:2405.13911Syntology official (archive's flag): 3 ran · 6 ran (of which 3 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 1 violated, 3 with no contract checked; 2 where Syntology's instrument failed) · 3 unverified (of 9 harvested samples) · 3 pointer-only (licence)
-
An Empirical Study and Analysis of Text-to-Image Generation Using Large Language Model-Powered Textual Representation 21 May 2024 · 1 repository · arXiv:2405.12914
-
Text-Video Retrieval with Global-Local Semantic Consistent Learning 21 May 2024 · 1 repository · arXiv:2405.12710
-
WorldAfford: Affordance Grounding based on Natural Language Instructions 21 May 2024 · 0 repositories · arXiv:2405.12461
-
Mammo-CLIP: A Vision Language Foundation Model to Enhance Data Efficiency and Robustness in Mammography 20 May 2024 · 1 repository · arXiv:2405.12255
-
Position-Guided Prompt Learning for Anomaly Detection in Chest X-Rays 20 May 2024 · 1 repository · arXiv:2405.11976
-
ColorFoil: Investigating Color Blindness in Large Vision and Language Models 19 May 2024 · 1 repository · arXiv:2405.11685
-
Hierarchical Selective Classification 19 May 2024 · 0 repositories · arXiv:2405.11533
-
Reproducibility Study of CDUL: CLIP-Driven Unsupervised Learning for Multi-Label Image Classification 19 May 2024 · 1 repository · arXiv:2405.11574
-
Track Anything Rapter(TAR) 19 May 2024 · 1 repository · arXiv:2405.11655
-
Unsupervised Image Prior via Prompt Learning and CLIP Semantic Guidance for Low-Light Image Enhancement 19 May 2024 · 0 repositories · arXiv:2405.11478
-
Enhancing Fine-Grained Image Classifications via Cascaded Vision Language Models 18 May 2024 · 0 repositories · arXiv:2405.11301
-
MediCLIP: Adapting CLIP for Few-shot Medical Image Anomaly Detection 18 May 2024 · 1 repository · arXiv:2405.11315
-
Revisiting the Robust Generalization of Adversarial Prompt Tuning 18 May 2024 · 0 repositories · arXiv:2405.11154
-
DiffAM: Diffusion-based Adversarial Makeup Transfer for Facial Privacy Protection 16 May 2024 · 2 repositories · arXiv:2405.09882Syntology official (archive's flag): 12 ran · 12 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 3 where Syntology's instrument failed) · 2 unverified (of 14 harvested samples) · 14 pointer-only (licence)
-
Harmonizing Generalization and Personalization in Federated Prompt Learning 16 May 2024 · 1 repository · arXiv:2405.09771Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 4 where Syntology's instrument failed) · 7 unverified (of 14 harvested samples) · 14 pointer-only (licence)
-
Natural Language Can Help Bridge the Sim2Real Gap 16 May 2024 · 0 repositories · arXiv:2405.10020
-
SHiNe: Semantic Hierarchy Nexus for Open-vocabulary Object Detection 16 May 2024 · 1 repository · arXiv:2405.10053
-
CLIP with Quality Captions: A Strong Pretraining for Vision Tasks 14 May 2024 · 0 repositories · arXiv:2405.08911
-
CLIP-Powered TASS: Target-Aware Single-Stream Network for Audio-Visual Question Answering 13 May 2024 · 0 repositories · arXiv:2405.07451
-
Investigating the Semantic Robustness of CLIP-based Zero-Shot Anomaly Segmentation 13 May 2024 · 0 repositories · arXiv:2405.07969
-
MoVL:Exploring Fusion Strategies for the Domain-Adaptive Application of Pretrained Models in Medical Imaging Tasks 13 May 2024 · 0 repositories · arXiv:2405.07411
-
Sakuga-42M Dataset: Scaling Up Cartoon Research 13 May 2024 · 0 repositories · arXiv:2405.07425
-
Zero Shot Context-Based Object Segmentation using SLIP (SAM+CLIP) 12 May 2024 · 1 repository · arXiv:2405.07284
-
Non-confusing Generation of Customized Concepts in Diffusion Models 11 May 2024 · 0 repositories · arXiv:2405.06914
-
RETTA: Retrieval-Enhanced Test-Time Adaptation for Zero-Shot Video Captioning 11 May 2024 · 0 repositories · arXiv:2405.07046
-
Decoding Emotions in Abstract Art: Cognitive Plausibility of CLIP in Recognizing Color-Emotion Associations 10 May 2024 · 0 repositories · arXiv:2405.06319
-
Enhancing Weakly Supervised Semantic Segmentation with Multi-modal Foundation Models: An End-to-End Approach 10 May 2024 · 0 repositories · arXiv:2405.06586
-
Open Challenges and Opportunities in Federated Foundation Models Towards Biomedical Healthcare 10 May 2024 · 0 repositories · arXiv:2405.06784
-
Enhanced Multimodal Content Moderation of Children's Videos using Audiovisual Fusion 9 May 2024 · 1 repository · arXiv:2405.06128
-
Exploring Text-Guided Single Image Editing for Remote Sensing Images 9 May 2024 · 1 repository · arXiv:2405.05769
-
Pre-trained Text-to-Image Diffusion Models Are Versatile Representation Learners for Control 9 May 2024 · 1 repository · arXiv:2405.05852Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
Attention-Driven Training-Free Efficiency Enhancement of Diffusion Models 8 May 2024 · 0 repositories · arXiv:2405.05252
-
Dual-Image Enhanced CLIP for Zero-Shot Anomaly Detection 8 May 2024 · 0 repositories · arXiv:2405.04782
-
OpenESS: Event-based Semantic Scene Understanding with Open Vocabularies 8 May 2024 · 1 repository · arXiv:2405.05259Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 2 where Syntology's instrument failed) · 2 unverified (of 9 harvested samples) · 9 pointer-only (licence)
-
Adapting Dual-encoder Vision-language Models for Paraphrased Retrieval 6 May 2024 · 0 repositories · arXiv:2405.03190
-
CICA: Content-Injected Contrastive Alignment for Zero-Shot Document Image Classification 6 May 2024 · 0 repositories · arXiv:2405.03660
-
Light-VQA+: A Video Quality Assessment Model for Exposure Correction with Vision-Language Guidance 6 May 2024 · 1 repository · arXiv:2405.03333
-
iSEARLE: Improving Textual Inversion for Zero-Shot Composed Image Retrieval 5 May 2024 · 2 repositories · arXiv:2405.02951Syntology official: no sample here; runs from other or unrecorded repositories · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 1 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Source-Free Domain Adaptation Guided by Vision and Vision-Language Pre-Training 5 May 2024 · 1 repository · arXiv:2405.02954
-
Enhancing Vision-Language Models Generalization via Diversity-Driven Novel Feature Synthesis 4 May 2024 · 0 repositories · arXiv:2405.02586
-
Improving Concept Alignment in Vision-Language Concept Bottleneck Models 3 May 2024 · 1 repository · arXiv:2405.01825
-
Multi-method Integration with Confidence-based Weighting for Zero-shot Image Classification 3 May 2024 · 0 repositories · arXiv:2405.02155
-
On the test-time zero-shot generalization of vision-language models: Do we really need prompt learning? 3 May 2024 · 1 repository · arXiv:2405.02266Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
EchoScene: Indoor Scene Generation via Information Echo over Scene Graph Diffusion 2 May 2024 · 1 repository · arXiv:2405.00915
-
Language-Enhanced Latent Representations for Out-of-Distribution Detection in Autonomous Driving 2 May 2024 · 0 repositories · arXiv:2405.01691
-
On Mechanistic Knowledge Localization in Text-to-Image Generative Models 2 May 2024 · 1 repository · arXiv:2405.01008Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Technical Report of NICE Challenge at CVPR 2024: Caption Re-ranking Evaluation Using Ensembled CLIP and Consensus Scores 2 May 2024 · 1 repository · arXiv:2405.01028
-
CLIPArTT: Adaptation of CLIP to New Domains at Test Time 1 May 2024 · 1 repository · arXiv:2405.00754Syntology official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 5 where Syntology's instrument failed) · 4 unverified (of 15 harvested samples) · 15 pointer-only (licence)
-
TexSliders: Diffusion-Based Texture Editing in CLIP Space 1 May 2024 · 0 repositories · arXiv:2405.00672
-
ESP-Zero: Unsupervised enhancement of zero-shot classification for Extremely Sparse Point cloud 30 Apr 2024 · 0 repositories · arXiv:2404.19639
-
Espresso: Robust Concept Filtering in Text-to-Image Models 30 Apr 2024 · 0 repositories · arXiv:2404.19227
-
MetaCoCo: A New Few-Shot Classification Benchmark with Spurious Correlation 30 Apr 2024 · 1 repository · arXiv:2404.19644Syntology official: harvested, nothing ran · 0 ran · 2 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Modeling Caption Diversity in Contrastive Vision-Language Pretraining 30 Apr 2024 · 1 repository · arXiv:2405.00740Syntology official (archive's flag): 14 ran · 14 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 1 honoured, 0 violated, 9 with no contract checked; 4 where Syntology's instrument failed) · 8 unverified (of 22 harvested samples) · 22 pointer-only (licence)
-
PEVA-Net: Prompt-Enhanced View Aggregation Network for Zero/Few-Shot Multi-View 3D Shape Recognition 30 Apr 2024 · 0 repositories · arXiv:2404.19168
-
Revisiting the Adversarial Robustness of Vision Language Models: a Multimodal Perspective 30 Apr 2024 · 1 repository · arXiv:2404.19287Syntology official (archive's flag): 18 ran · 18 ran (of which 0 constructed an object rather than computing a result; 12 with no instrument failure: 1 honoured, 0 violated, 11 with no contract checked; 6 where Syntology's instrument failed) · 4 unverified (of 22 harvested samples) · 5 pointer-only (licence)
-
Weighted Point Cloud Embedding for Multimodal Contrastive Learning Toward Optimal Similarity Metric 30 Apr 2024 · 0 repositories · arXiv:2404.19228Syntology 5 ran (of which 4 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 7 harvested samples)
-
Breaking Through the Noisy Correspondence: A Robust Model for Image-Text Matching 29 Apr 2024 · 0 repositories
-
Dual-Modal Prompting for Sketch-Based Image Retrieval 29 Apr 2024 · 0 repositories · arXiv:2404.18695
-
Q-GroundCAM: Quantifying Grounding in Vision Language Models via GradCAM 29 Apr 2024 · 0 repositories · arXiv:2404.19128
-
Saliency Suppressed, Semantics Surfaced: Visual Transformations in Neural Networks and the Brain 29 Apr 2024 · 1 repository · arXiv:2404.18772
-
Spatio-Temporal Side Tuning Pre-trained Foundation Models for Video-based Pedestrian Attribute Recognition 27 Apr 2024 · 3 repositories · arXiv:2404.17929
-
FashionSD-X: Multimodal Fashion Garment Synthesis using Latent Diffusion 26 Apr 2024 · 0 repositories · arXiv:2404.18591
-
HYPE: Hyperbolic Entailment Filtering for Underspecified Images and Texts 26 Apr 2024 · 1 repository · arXiv:2404.17507
-
Learning text-to-video retrieval from image captioning 26 Apr 2024 · 0 repositories · arXiv:2404.17498
-
Open-Set Video-based Facial Expression Recognition with Human Expression-sensitive Prompting 26 Apr 2024 · 0 repositories · arXiv:2404.17100
-
Trinity Detector:text-assisted and attention mechanisms based spectral fusion for diffusion generation image detection 26 Apr 2024 · 0 repositories · arXiv:2404.17254
-
Learning Discriminative Spatio-temporal Representations for Semi-supervised Action Recognition 25 Apr 2024 · 0 repositories · arXiv:2404.16416
-
Revisiting Relevance Feedback for CLIP-based Interactive Image Retrieval 25 Apr 2024 · 0 repositories · arXiv:2404.16398