Methods › Computer Vision › Vision and Language Pre-Trained Models › CLIP › Papers, page 12
Contrastive Language-Image Pre-training
CLIP
Papers archive 2025-07-28
archive papers tagged: 3,094 · with a code link: 1,617 · where Syntology ran a sample: 649 (554 with a run with no instrument failure, 95 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (649 of 3,094 tagged: 554 with a run with no instrument failure, 95 where every run was a failure of Syntology's instrument)
Page 12 of 31: papers 1,101 to 1,200 of 3,094, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Distilling Vision-Language Foundation Models: A Data-Free Approach via Prompt Diversification 21 Jul 2024 · 0 repositories · arXiv:2407.15155
-
Prior Knowledge Integration via LLM Encoding and Pseudo Event Regulation for Video Moment Retrieval 21 Jul 2024 · 1 repository · arXiv:2407.15051Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 1 violated, 4 with no contract checked; 3 where Syntology's instrument failed) · 3 unverified (of 11 harvested samples) · 9 pointer-only (licence)
-
Rethinking Domain Adaptation and Generalization in the Era of CLIP 21 Jul 2024 · 0 repositories · arXiv:2407.15173
-
Adapt2Reward: Adapting Video-Language Models to Generalizable Robotic Rewards via Failure Prompts 20 Jul 2024 · 0 repositories · arXiv:2407.14872
-
Sim-CLIP: Unsupervised Siamese Adversarial Fine-Tuning for Robust and Semantically-Rich Vision-Language Models 20 Jul 2024 · 0 repositories · arXiv:2407.14971
-
A Benchmark for Gaussian Splatting Compression and Quality Assessment Study 19 Jul 2024 · 1 repository · arXiv:2407.14197
-
Braille-to-Speech Generator: Audio Generation Based on Joint Fine-Tuning of CLIP and Fastspeech2 19 Jul 2024 · 0 repositories · arXiv:2407.14212
-
Class-Incremental Learning with CLIP: Adaptive Representation Adjustment and Parameter Fusion 19 Jul 2024 · 1 repository · arXiv:2407.14143Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
Discover-then-Name: Task-Agnostic Concept Bottlenecks via Automated Concept Discovery 19 Jul 2024 · 1 repository · arXiv:2407.14499
-
HOTS3D: Hyper-Spherical Optimal Transport for Semantic Alignment of Text-to-3D Generation 19 Jul 2024 · 0 repositories · arXiv:2407.14419
-
Rethinking Visual Content Refinement in Low-Shot CLIP Adaptation 19 Jul 2024 · 1 repository · arXiv:2407.14117
-
CoAPT: Context Attribute words for Prompt Tuning 18 Jul 2024 · 0 repositories · arXiv:2407.13808
-
HazeCLIP: Towards Language Guided Real-World Image Dehazing 18 Jul 2024 · 1 repository · arXiv:2407.13719
-
New Capability to Look Up an ASL Sign from a Video Example 18 Jul 2024 · 0 repositories · arXiv:2407.13571
-
Robust Calibration of Large Vision-Language Adapters 18 Jul 2024 · 1 repository · arXiv:2407.13588Syntology official (archive's flag): 1 ran · 1 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified; the one sample that ran constructed an object rather than computing a result (of 1 harvested sample)
-
ClearCLIP: Decomposing CLIP Representations for Dense Vision-Language Inference 17 Jul 2024 · 0 repositories · arXiv:2407.12442
-
Direct Unlearning Optimization for Robust and Safe Text-to-Image Models 17 Jul 2024 · 0 repositories · arXiv:2407.21035
-
IMAGDressing-v1: Customizable Virtual Dressing 17 Jul 2024 · 1 repository · arXiv:2407.12705Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 1 pointer-only (licence)
-
JointDreamer: Ensuring Geometry Consistency and Text Congruence in Text-to-3D Generation via Joint Score Distillation 17 Jul 2024 · 0 repositories · arXiv:2407.12291
-
ModalChorus: Visual Probing and Alignment of Multi-modal Embeddings via Modal Fusion Map 17 Jul 2024 · 1 repository · arXiv:2407.12315
-
VCP-CLIP: A visual context prompting model for zero-shot anomaly segmentation 17 Jul 2024 · 1 repository · arXiv:2407.12276
-
VEON: Vocabulary-Enhanced Occupancy Prediction 17 Jul 2024 · 0 repositories · arXiv:2407.12294
-
An AI System for Continuous Knee Osteoarthritis Severity Grading Using Self-Supervised Anomaly Detection with Limited Data 16 Jul 2024 · 1 repository · arXiv:2407.11500
-
Continuous Embedding Attacks via Clipped Inputs in Jailbreaking Large Language Models 16 Jul 2024 · 1 repository · arXiv:2407.13796
-
LaMI-DETR: Open-Vocabulary Detection with Language Model Instruction 16 Jul 2024 · 1 repository · arXiv:2407.11335Syntology official (archive's flag): 3 ran · 3 ran (of which 2 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples)
-
Large Visual-Language Models Are Also Good Classifiers: A Study of In-Context Multimodal Fake News Detection 16 Jul 2024 · 0 repositories · arXiv:2407.12879
-
Unlearning Targeted Information via Single Layer Unlearning Gradient 16 Jul 2024 · 1 repository · arXiv:2407.11867Syntology official (archive's flag): 16 ran · 16 ran (of which 0 constructed an object rather than computing a result; 12 with no instrument failure: 0 honoured, 0 violated, 12 with no contract checked; 4 where Syntology's instrument failed) · 8 unverified (of 24 harvested samples) · 24 pointer-only (licence)
-
Accessing Vision Foundation Models at ImageNet-level Costs 15 Jul 2024 · 1 repository · arXiv:2407.10366Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 1 honoured, 0 violated, 2 with no contract checked; 1 where Syntology's instrument failed) · 6 unverified (of 10 harvested samples)
-
DataDream: Few-shot Guided Dataset Generation 15 Jul 2024 · 2 repositories · arXiv:2407.10910Syntology official (archive's flag): 17 ran · 17 ran (of which 0 constructed an object rather than computing a result; 16 with no instrument failure: 1 honoured, 0 violated, 15 with no contract checked; 1 where Syntology's instrument failed) · 4 unverified (of 21 harvested samples) · 21 pointer-only (licence)
-
How and where does CLIP process negation? 15 Jul 2024 · 0 repositories · arXiv:2407.10488
-
Quantized Prompt for Efficient Generalization of Vision-Language Models 15 Jul 2024 · 1 repository · arXiv:2407.10704
-
Unconstrained Open Vocabulary Image Classification: Zero-Shot Transfer from Text to Image via CLIP Inversion 15 Jul 2024 · 2 repositories · arXiv:2407.11211
-
CLIP-Guided Generative Networks for Transferable Targeted Adversarial Attacks 14 Jul 2024 · 1 repository · arXiv:2407.10179Syntology official (archive's flag): 2 ran · 3 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 2 where Syntology's instrument failed) · 2 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
Dense Multimodal Alignment for Open-Vocabulary 3D Scene Understanding 13 Jul 2024 · 0 repositories · arXiv:2407.09781
-
LAPT: Label-driven Automated Prompt Tuning for OOD Detection with Vision-Language Models 12 Jul 2024 · 2 repositories · arXiv:2407.08966Syntology community repositories only · 15 ran (of which 0 constructed an object rather than computing a result; 11 with no instrument failure: 0 honoured, 2 violated, 9 with no contract checked; 4 where Syntology's instrument failed) · 2 unverified (of 17 harvested samples) · 4 pointer-only (licence)
-
Open Vocabulary Multi-Label Video Classification 12 Jul 2024 · 0 repositories · arXiv:2407.09073
-
Surgical Text-to-Image Generation 12 Jul 2024 · 0 repositories · arXiv:2407.09230
-
Emergent Visual-Semantic Hierarchies in Image-Text Representations 11 Jul 2024 · 1 repository · arXiv:2407.08521Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Enhancing Robustness of Vision-Language Models through Orthogonality Learning and Self-Regularization 11 Jul 2024 · 0 repositories · arXiv:2407.08374
-
Explore the Potential of CLIP for Training-Free Open Vocabulary Semantic Segmentation 11 Jul 2024 · 1 repository · arXiv:2407.08268Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 5 harvested samples)
-
Fine-Tuning Stable Diffusion XL for Stylistic Icon Generation: A Comparison of Caption Size 11 Jul 2024 · 0 repositories · arXiv:2407.08513
-
LDRE: LLM-based Divergent Reasoning and Ensemble for Zero-Shot Composed Image Retrieval 11 Jul 2024 · 2 repositories
-
CosmoCLIP: Generalizing Large Vision-Language Models for Astronomical Imaging 10 Jul 2024 · 0 repositories · arXiv:2407.07315
-
Unified Embedding Alignment for Open-Vocabulary Video Instance Segmentation 10 Jul 2024 · 1 repository · arXiv:2407.07427
-
Video In-context Learning 10 Jul 2024 · 0 repositories · arXiv:2407.07356
-
Zero-Shot Class Unlearning in CLIP with Synthetic Samples 10 Jul 2024 · 1 repository · arXiv:2407.07485
-
CEIA: CLIP-Based Event-Image Alignment for Open-World Event-Based Understanding 9 Jul 2024 · 0 repositories · arXiv:2407.06611
-
CoLA: Conditional Dropout and Language-driven Robust Dual-modal Salient Object Detection 9 Jul 2024 · 1 repository · arXiv:2407.06780Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 5 where Syntology's instrument failed) · 3 unverified (of 11 harvested samples) · 4 pointer-only (licence)
-
Exploring Scalability of Self-Training for Open-Vocabulary Temporal Action Localization 9 Jul 2024 · 1 repository · arXiv:2407.07024
-
Fine-Tuning Attention Modules Only: Enhancing Weight Disentanglement in Task Arithmetic 9 Jul 2024 · 2 repositories · arXiv:2407.07089Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 2 unverified (of 5 harvested samples)
-
Graph-Based Captioning: Enhancing Visual Descriptions by Interconnecting Region Captions 9 Jul 2024 · 0 repositories · arXiv:2407.06723
-
Contrastive Learning of Preferences with a Contextual InfoNCE Loss 8 Jul 2024 · 0 repositories · arXiv:2407.05898
-
Deciphering the Role of Representation Disentanglement: Investigating Compositional Generalization in CLIP Models 8 Jul 2024 · 1 repository · arXiv:2407.05897
-
FALIP: Visual Prompt as Foveal Attention Boosts CLIP Zero-Shot Performance 8 Jul 2024 · 0 repositories · arXiv:2407.05578
-
Learning to Adapt Category Consistent Meta-Feature of CLIP for Few-Shot Classification 8 Jul 2024 · 0 repositories · arXiv:2407.05647
-
Leveraging Transformers for Weakly Supervised Object Localization in Unconstrained Videos 8 Jul 2024 · 1 repository · arXiv:2407.06018
-
Pseudo-triplet Guided Few-shot Composed Image Retrieval 8 Jul 2024 · 0 repositories · arXiv:2407.06001
-
Towards Bridging the Cross-modal Semantic Gap for Multi-modal Recommendation 7 Jul 2024 · 1 repository · arXiv:2407.05420
-
The Solution for Language-Enhanced Image New Category Discovery 6 Jul 2024 · 0 repositories · arXiv:2407.04994
-
The Solution for the 5th GCAIAC Zero-shot Referring Expression Comprehension Challenge 6 Jul 2024 · 0 repositories · arXiv:2407.04998
-
AWT: Transferring Vision-Language Models via Augmentation, Weighting, and Transportation 5 Jul 2024 · 1 repository · arXiv:2407.04603Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Elevating All Zero-Shot Sketch-Based Image Retrieval Through Multimodal Prompt Learning 5 Jul 2024 · 1 repository · arXiv:2407.04207
-
CLIP-DR: Textual Knowledge-Guided Diabetic Retinopathy Grading with Ranking-aware Prompting 4 Jul 2024 · 1 repository · arXiv:2407.04068
-
Do Generalised Classifiers really work on Human Drawn Sketches? 4 Jul 2024 · 1 repository · arXiv:2407.03893
-
EMPL: A novel Efficient Meta Prompt Learning Framework for Few-shot Unsupervised Domain Adaptation 4 Jul 2024 · 0 repositories · arXiv:2407.04066
-
MRIR: Integrating Multimodal Insights for Diffusion-based Realistic Image Restoration 4 Jul 2024 · 0 repositories · arXiv:2407.03635
-
SOWA: Adapting Hierarchical Frozen Window Self-Attention to Visual-Language Models for Better Anomaly Detection 4 Jul 2024 · 1 repository · arXiv:2407.03634
-
SAFT: Towards Out-of-Distribution Generalization in Fine-Tuning 3 Jul 2024 · 0 repositories · arXiv:2407.03036
-
Boosting Consistency in Story Visualization with Rich-Contextual Conditional Diffusion Models 2 Jul 2024 · 1 repository · arXiv:2407.02482Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 2 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
FineCLIPER: Multi-modal Fine-grained CLIP for Dynamic Facial Expression Recognition with AdaptERs 2 Jul 2024 · 0 repositories · arXiv:2407.02157
-
Lung-CADex: Fully automatic Zero-Shot Detection and Classification of Lung Nodules in Thoracic CT Images 2 Jul 2024 · 0 repositories · arXiv:2407.02625
-
Magic Insert: Style-Aware Drag-and-Drop 2 Jul 2024 · 0 repositories · arXiv:2407.02489
-
CLIP the Divergence: Language-guided Unsupervised Domain Adaptation 1 Jul 2024 · 0 repositories · arXiv:2407.01842
-
Fast and Efficient: Mask Neural Fields for 3D Scene Segmentation 1 Jul 2024 · 1 repository · arXiv:2407.01220
-
FastCLIP: A Suite of Optimization Techniques to Accelerate CLIP Training with Limited Resources 1 Jul 2024 · 1 repository · arXiv:2407.01445
-
GalLoP: Learning Global and Local Prompts for Vision-Language Models 1 Jul 2024 · 1 repository · arXiv:2407.01400Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples)
-
Learning Robust 3D Representation from CLIP via Dual Denoising 1 Jul 2024 · 0 repositories · arXiv:2407.00905
-
Semantic Compositions Enhance Vision-Language Contrastive Learning 1 Jul 2024 · 0 repositories · arXiv:2407.01408
-
SignCLIP: Connecting Text and Sign Language by Contrastive Learning 1 Jul 2024 · 1 repository · arXiv:2407.01264
-
Unveiling Glitches: A Deep Dive into Image Encoding Bugs within CLIP 30 Jun 2024 · 0 repositories · arXiv:2407.00592
-
EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model 28 Jun 2024 · 1 repository · arXiv:2406.20076Syntology official (archive's flag): 1 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 2 pointer-only (licence)
-
GM-DF: Generalized Multi-Scenario Deepfake Detection 28 Jun 2024 · 1 repository · arXiv:2406.20078
-
PathGen-1.6M: 1.6 Million Pathology Image-text Pairs Generation through Multi-agent Collaboration 28 Jun 2024 · 1 repository · arXiv:2407.00203
-
A Sanity Check for AI-generated Image Detection 27 Jun 2024 · 2 repositories · arXiv:2406.19435Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 8 harvested samples) · 3 pointer-only (licence)
-
CLIP3D-AD: Extending CLIP for 3D Few-Shot Anomaly Detection with Multi-View Images Generation 27 Jun 2024 · 0 repositories · arXiv:2406.18941
-
Rethinking and Defending Protective Perturbation in Personalized Diffusion Models 27 Jun 2024 · 1 repository · arXiv:2406.18944
-
3D Feature Distillation with Object-Centric Priors 26 Jun 2024 · 0 repositories · arXiv:2406.18742
-
BioTrove: A Large Curated Image Dataset Enabling AI for Biodiversity 25 Jun 2024 · 2 repositories · arXiv:2406.17720Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
ET tu, CLIP? Addressing Common Object Errors for Unseen Environments 25 Jun 2024 · 0 repositories · arXiv:2406.17876
-
Mitigate the Gap: Investigating Approaches for Improving Cross-Modal Alignment in CLIP 25 Jun 2024 · 1 repository · arXiv:2406.17639Syntology official (archive's flag): 16 ran · 16 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 3 violated, 7 with no contract checked; 6 where Syntology's instrument failed) · 6 unverified (of 22 harvested samples) · 22 pointer-only (licence)
-
InterCLIP-MEP: Interactive CLIP and Memory-Enhanced Predictor for Multi-modal Sarcasm Detection 24 Jun 2024 · 1 repository · arXiv:2406.16464
-
Video-Infinity: Distributed Long Video Generation 24 Jun 2024 · 0 repositories · arXiv:2406.16260
-
Vision-Language Consistency Guided Multi-modal Prompt Learning for Blind AI Generated Image Quality Assessment 24 Jun 2024 · 1 repository · arXiv:2406.16641
-
Multi-Scale Temporal Difference Transformer for Video-Text Retrieval 23 Jun 2024 · 0 repositories · arXiv:2406.16111
-
CLIP-Decoder : ZeroShot Multilabel Classification using Multimodal CLIP Aligned Representation 21 Jun 2024 · 1 repository · arXiv:2406.14830Syntology 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Improving Interpretability and Robustness for the Detection of AI-Generated Images 21 Jun 2024 · 0 repositories · arXiv:2406.15035
-
African or European Swallow? Benchmarking Large Vision-Language Models for Fine-Grained Object Classification 20 Jun 2024 · 1 repository · arXiv:2406.14496Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 9 harvested samples)
-
ARDuP: Active Region Video Diffusion for Universal Policies 19 Jun 2024 · 0 repositories · arXiv:2406.13301
-
CLIP-Branches: Interactive Fine-Tuning for Text-Image Retrieval 19 Jun 2024 · 1 repository · arXiv:2406.13322
-
IntCoOp: Interpretability-Aware Vision-Language Prompt Tuning 19 Jun 2024 · 0 repositories · arXiv:2406.13683