Methods › Computer Vision › Vision and Language Pre-Trained Models › CLIP › Papers, page 2
Contrastive Language-Image Pre-training
CLIP
Papers archive 2025-07-28
archive papers tagged: 3,094 · with a code link: 1,617 · where Syntology ran a sample: 649 (554 with a run with no instrument failure, 95 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (649 of 3,094 tagged: 554 with a run with no instrument failure, 95 where every run was a failure of Syntology's instrument)
Page 2 of 31: papers 101 to 200 of 3,094, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
EvdCLIP: Improving Vision-Language Retrieval with Entity Visual Descriptions from Large Language Models 24 May 2025 · 0 repositories · arXiv:2505.18594
-
LORE: Lagrangian-Optimized Robust Embeddings for Visual Encoders 24 May 2025 · 1 repository · arXiv:2505.18884
-
REGen: Multimodal Retrieval-Embedded Generation for Long-to-Short Video Editing 24 May 2025 · 0 repositories · arXiv:2505.18880
-
Self-Supervised and Generalizable Tokenization for CLIP-Based 3D Understanding 24 May 2025 · 0 repositories · arXiv:2505.18819
-
TNG-CLIP:Training-Time Negation Data Generation for Negation Awareness of CLIP 24 May 2025 · 0 repositories · arXiv:2505.18434
-
Clip4Retrofit: Enabling Real-Time Image Labeling on Edge Devices via Cross-Architecture CLIP Distillation 23 May 2025 · 0 repositories · arXiv:2505.18039
-
From Flight to Insight: Semantic 3D Reconstruction for Aerial Inspection via Gaussian Splatting and Language-Guided Segmentation 23 May 2025 · 0 repositories · arXiv:2505.17402
-
ICPL-ReID: Identity-Conditional Prompt Learning for Multi-Spectral Object Re-Identification 23 May 2025 · 1 repository · arXiv:2505.17821
-
DeCafNet: Delegate and Conquer for Efficient Temporal Grounding in Long Videos 22 May 2025 · 1 repository · arXiv:2505.16376
-
Investigating Fine- and Coarse-grained Structural Correspondences Between Deep Neural Networks and Human Object Image Similarity Judgments Using Unsupervised Alignment 22 May 2025 · 0 repositories · arXiv:2505.16419
-
Explainable embeddings with Distance Explainer 21 May 2025 · 0 repositories · arXiv:2505.15516
-
Few-Shot Adversarial Low-Rank Fine-Tuning of Vision-Language Models 21 May 2025 · 0 repositories · arXiv:2505.15130
-
Image-to-Image Translation with Diffusion Transformers and CLIP-Based Image Conditioning 21 May 2025 · 0 repositories · arXiv:2505.16001
-
MoRE-Brain: Routed Mixture of Experts for Interpretable and Generalizable Cross-Subject fMRI Visual Decoding 21 May 2025 · 0 repositories · arXiv:2505.15946Syntology 2 ran (of which 2 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified; every one of the 2 samples that ran constructed an object rather than computing a result (of 4 harvested samples) · 4 pointer-only (licence)
-
Multimodal Conditional Information Bottleneck for Generalizable AI-Generated Image Detection 21 May 2025 · 1 repository · arXiv:2505.15217
-
Prompt Tuning Vision Language Models with Margin Regularizer for Few-Shot Learning under Distribution Shifts 21 May 2025 · 1 repository · arXiv:2505.15506
-
TAGS: 3D Tumor-Adaptive Guidance for SAM 21 May 2025 · 0 repositories · arXiv:2505.17096
-
Beginning with You: Perceptual-Initialization Improves Vision-Language Representation and Alignment 20 May 2025 · 0 repositories · arXiv:2505.14204
-
Breaking Language Barriers or Reinforcing Bias? A Study of Gender and Racial Disparities in Multilingual Contrastive Vision Language Models 20 May 2025 · 0 repositories · arXiv:2505.14160
-
ReactDiff: Latent Diffusion for Facial Reaction Generation 20 May 2025 · 1 repository · arXiv:2505.14151
-
From Local Details to Global Context: Advancing Vision-Language Models with Attention-Based Selection 19 May 2025 · 1 repository · arXiv:2505.13233
-
ReSW-VL: Representation Learning for Surgical Workflow Analysis Using Vision-Language Model 19 May 2025 · 0 repositories · arXiv:2505.13746
-
SPKLIP: Aligning Spike Video Streams with Natural Language 19 May 2025 · 0 repositories · arXiv:2505.12656
-
StarFT: Robust Fine-tuning of Zero-shot Models via Spuriosity Alignment 19 May 2025 · 1 repository · arXiv:2505.13232
-
Uniformity First: Uniformity-aware Test-time Adaptation of Vision-language Models against Image Corruption 19 May 2025 · 1 repository · arXiv:2505.12912
-
CLIP-aware Domain-Adaptive Super-Resolution 18 May 2025 · 0 repositories · arXiv:2505.12391
-
CPGD: Toward Stable Rule-based Reinforcement Learning for Language Models 18 May 2025 · 1 repository · arXiv:2505.12504
-
Guiding Diffusion with Deep Geometric Moments: Balancing Fidelity and Variation 18 May 2025 · 0 repositories · arXiv:2505.12486
-
Imagination-Limited Q-Learning for Offline Reinforcement Learning 18 May 2025 · 0 repositories · arXiv:2505.12211
-
KGAlign: Joint Semantic-Structural Knowledge Encoding for Multimodal Fake News Detection 18 May 2025 · 1 repository · arXiv:2505.14714
-
Video-GPT via Next Clip Diffusion 18 May 2025 · 1 repository · arXiv:2505.12489
-
ViEEG: Hierarchical Neural Coding with Cross-Modal Progressive Enhancement for EEG-Based Visual Decoding 18 May 2025 · 0 repositories · arXiv:2505.12408
-
GenZSL: Generative Zero-Shot Learning Via Inductive Variational Autoencoder 17 May 2025 · 1 repository · arXiv:2505.11882Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
MoCLIP: Motion-Aware Fine-Tuning and Distillation of CLIP for Human Motion Generation 16 May 2025 · 0 repositories · arXiv:2505.10810
-
Concept Drift Guided LayerNorm Tuning for Efficient Multimodal Metaphor Identification 16 May 2025 · 2 repositories · arXiv:2505.11237
-
Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation 16 May 2025 · 1 repository · arXiv:2505.11383
-
Patho-R1: A Multimodal Reinforcement Learning-Based Pathology Expert Reasoner 16 May 2025 · 1 repository · arXiv:2505.11404
-
DPSeg: Dual-Prompt Cost Volume Learning for Open-Vocabulary Semantic Segmentation 16 May 2025 · 0 repositories · arXiv:2505.11676
-
Open Set Domain Adaptation with Vision-language models via Gradient-aware Separation 16 May 2025 · 0 repositories · arXiv:2505.13507
-
Semantically-Aware Game Image Quality Assessment 16 May 2025 · 0 repositories · arXiv:2505.11724
-
CLIP Embeddings for AI-Generated Image Detection: A Few-Shot Study with Lightweight Classifier 15 May 2025 · 0 repositories · arXiv:2505.10664
-
AdaptCLIP: Adapting CLIP for Universal Visual Anomaly Detection 15 May 2025 · 1 repository · arXiv:2505.09926
-
Does Feasibility Matter? Understanding the Impact of Feasibility on Synthetic Training Data 15 May 2025 · 1 repository · arXiv:2505.10551
-
MMRL++: Parameter-Efficient and Interaction-Aware Representation Learning for Vision-Language Models 15 May 2025 · 1 repository · arXiv:2505.10088
-
MSCI: Addressing CLIP's Inherent Limitations for Compositional Zero-Shot Learning 15 May 2025 · 1 repository · arXiv:2505.10289Syntology official (archive's flag): 22 ran · 23 ran (of which 13 constructed an object rather than computing a result; 19 with no instrument failure: 0 honoured, 0 violated, 19 with no contract checked; 4 where Syntology's instrument failed) · 5 unverified (of 28 harvested samples) · 28 pointer-only (licence)
-
BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset 14 May 2025 · 1 repository · arXiv:2505.09568Syntology official (archive's flag): 1 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Denoising and Alignment: Rethinking Domain Generalization for Multimodal Face Anti-Spoofing 14 May 2025 · 0 repositories · arXiv:2505.09484
-
VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models 13 May 2025 · 1 repository · arXiv:2505.08455
-
Visually Guided Decoding: Gradient-Free Hard Prompt Inversion with Language Models 13 May 2025 · 1 repository · arXiv:2505.08622Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples)
-
Addressing degeneracies in latent interpolation for diffusion models 12 May 2025 · 0 repositories · arXiv:2505.07481
-
Beyond CLIP Generalization: Against Forward&Backward Forgetting Adapter for Continual Learning of Vision-Language Models 12 May 2025 · 0 repositories · arXiv:2505.07690
-
DanceGRPO: Unleashing GRPO on Visual Generation 12 May 2025 · 1 repository · arXiv:2505.07818Syntology 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 1 honoured, 1 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples)
-
SLAG: Scalable Language-Augmented Gaussian Splatting 12 May 2025 · 0 repositories · arXiv:2505.08124
-
Replay-Based Continual Learning with Dual-Layered Distillation and a Streamlined U-Net for Efficient Text-to-Image Generation 11 May 2025 · 0 repositories · arXiv:2505.06995
-
Whitened CLIP as a Likelihood Surrogate of Images and Captions 11 May 2025 · 1 repository · arXiv:2505.06934Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
HCMA: Hierarchical Cross-model Alignment for Grounded Text-to-Image Generation 10 May 2025 · 1 repository · arXiv:2505.06512
-
METOR: A Unified Framework for Mutual Enhancement of Objects and Relationships in Open-vocabulary Video Visual Relationship Detection 10 May 2025 · 1 repository · arXiv:2505.06663
-
Model Steering: Learning with a Reference Model Improves Generalization Bounds and Scaling Laws 10 May 2025 · 1 repository · arXiv:2505.06699
-
PromptIQ: Who Cares About Prompts? Let System Handle It -- A Component-Aware Framework for T2I Generation 9 May 2025 · 0 repositories · arXiv:2505.06467
-
Task-Adapter++: Task-specific Adaptation with Order-aware Alignment for Few-shot Action Recognition 9 May 2025 · 1 repository · arXiv:2505.06002
-
Does CLIP perceive art the same way we do? 8 May 2025 · 0 repositories · arXiv:2505.05229
-
FG-CLIP: Fine-Grained Visual and Textual Alignment 8 May 2025 · 1 repository · arXiv:2505.05071Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Hearing and Seeing Through CLIP: A Framework for Self-Supervised Sound Source Localization 8 May 2025 · 1 repository · arXiv:2505.05343
-
In-Context Learning for Label-Efficient Cancer Image Classification in Oncology 8 May 2025 · 0 repositories · arXiv:2505.08798
-
OpenworldAUC: Towards Unified Evaluation and Optimization for Open-world Prompt Tuning 8 May 2025 · 1 repository · arXiv:2505.05180
-
PIDiff: Image Customization for Personalized Identities with Diffusion Models 8 May 2025 · 0 repositories · arXiv:2505.05081
-
Probabilistic Embeddings for Frozen Vision-Language Models: Uncertainty Quantification with Gaussian Process Latent Variable Models 8 May 2025 · 1 repository · arXiv:2505.05163
-
Split Matching for Inductive Zero-shot Semantic Segmentation 8 May 2025 · 0 repositories · arXiv:2505.05023
-
ULFine: Unbiased Lightweight Fine-tuning for Foundation-Model-Assisted Long-Tailed Semi-Supervised Learning 8 May 2025 · 0 repositories · arXiv:2505.05062
-
X-Transfer Attacks: Towards Super Transferable Adversarial Attacks on CLIP 8 May 2025 · 1 repository · arXiv:2505.05528Syntology official (archive's flag): 5 ran · 5 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 4 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples)
-
DeCLIP: Decoupled Learning for Open-Vocabulary Dense Perception 7 May 2025 · 1 repository · arXiv:2505.04410Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample)
-
OpenVision: A Fully-Open, Cost-Effective Family of Advanced Vision Encoders for Multimodal Learning 7 May 2025 · 0 repositories · arXiv:2505.04601
-
Replay to Remember (R2R): An Efficient Uncertainty-driven Unsupervised Continual Learning Framework Using Generative Replay 7 May 2025 · 0 repositories · arXiv:2505.04787
-
A Vision-Language Model for Focal Liver Lesion Classification 6 May 2025 · 0 repositories · arXiv:2505.03350
-
Multimodal Benchmarking and Recommendation of Text-to-Image Generation Models 6 May 2025 · 1 repository · arXiv:2505.04650
-
Panoramic Out-of-Distribution Segmentation 6 May 2025 · 1 repository · arXiv:2505.03539
-
Recent Advances in Out-of-Distribution Detection with CLIP-Like Models: A Survey 5 May 2025 · 0 repositories · arXiv:2505.02448
-
TeDA: Boosting Vision-Lanuage Models for Zero-Shot 3D Object Retrieval via Testing-time Distribution Alignment 5 May 2025 · 1 repository · arXiv:2505.02325
-
Using Knowledge Graphs to harvest datasets for efficient CLIP model training 5 May 2025 · 1 repository · arXiv:2505.02746
-
Compositional Image-Text Matching and Retrieval by Grounding Entities 4 May 2025 · 0 repositories · arXiv:2505.02278
-
Robust AI-Generated Face Detection with Imbalanced Data 4 May 2025 · 1 repository · arXiv:2505.02182
-
TxP: Reciprocal Generation of Ground Pressure Dynamics and Activity Descriptions for Improving Human Activity Recognition 4 May 2025 · 1 repository · arXiv:2505.02052
-
Topology-Aware CLIP Few-Shot Learning 3 May 2025 · 0 repositories · arXiv:2505.01694
-
Carbon Aware Transformers Through Joint Model-Hardware Optimization 2 May 2025 · 1 repository · arXiv:2505.01386
-
Efficient Vocabulary-Free Fine-Grained Visual Recognition in the Age of Multimodal LLMs 2 May 2025 · 0 repositories · arXiv:2505.01064
-
FlowDubber: Movie Dubbing with LLM-based Semantic-aware Learning and Flow Matching based Voice Enhancing 2 May 2025 · 0 repositories · arXiv:2505.01263
-
Model See Model Do: Speech-Driven Facial Animation with Style Control 2 May 2025 · 0 repositories · arXiv:2505.01319
-
VSC: Visual Search Compositional Text-to-Image Diffusion Model 2 May 2025 · 0 repositories · arXiv:2505.01104
-
AnimalMotionCLIP: Embedding motion in CLIP for Animal Behavior Analysis 30 Apr 2025 · 0 repositories · arXiv:2505.00569
-
Vision-Language Model-Based Semantic-Guided Imaging Biomarker for Early Lung Cancer Detection 30 Apr 2025 · 0 repositories · arXiv:2504.21344
-
ClearVision: Leveraging CycleGAN and SigLIP-2 for Robust All-Weather Classification in Traffic Camera Imagery 28 Apr 2025 · 0 repositories · arXiv:2504.19684
-
A Review of 3D Object Detection with Vision-Language Models 25 Apr 2025 · 0 repositories · arXiv:2504.18738
-
CLIPSE -- a minimalistic CLIP-based image search engine for research 24 Apr 2025 · 1 repository · arXiv:2504.17643
-
A Survey of Foundation Model-Powered Recommender Systems: From Feature-Based, Generative to Agentic Paradigms 23 Apr 2025 · 0 repositories · arXiv:2504.16420
-
DP2FL: Dual Prompt Personalized Federated Learning in Foundation Models 23 Apr 2025 · 0 repositories · arXiv:2504.16357
-
FrogDogNet: Fourier frequency Retained visual prompt Output Guidance for Domain Generalization of CLIP in Remote Sensing 23 Apr 2025 · 0 repositories · arXiv:2504.16433
-
Backdoor Defense in Diffusion Models via Spatial Attention Unlearning 21 Apr 2025 · 0 repositories · arXiv:2504.18563
-
GenCLIP: Generalizing CLIP Prompts for Zero-shot Anomaly Detection 21 Apr 2025 · 0 repositories · arXiv:2504.14919
-
Hierarchical Attention Fusion of Visual and Textual Representations for Cross-Domain Sequential Recommendation 21 Apr 2025 · 0 repositories · arXiv:2504.15085
-
Text-to-Decision Agent: Learning Generalist Policies from Natural Language Supervision 21 Apr 2025 · 0 repositories · arXiv:2504.15046