Methods › Computer Vision › Vision and Language Pre-Trained Models › CLIP › Papers, page 4
Contrastive Language-Image Pre-training
CLIP
Papers archive 2025-07-28
archive papers tagged: 3,094 · with a code link: 1,617 · where Syntology ran a sample: 649 (554 with a run with no instrument failure, 95 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (649 of 3,094 tagged: 554 with a run with no instrument failure, 95 where every run was a failure of Syntology's instrument)
Page 4 of 31: papers 301 to 400 of 3,094, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
SPNeRF: Open Vocabulary 3D Neural Scene Segmentation with Superpoints 19 Mar 2025 · 0 repositories · arXiv:2503.15712
-
TULIP: Towards Unified Language-Image Pretraining 19 Mar 2025 · 0 repositories · arXiv:2503.15485
-
DAPO: An Open-Source LLM Reinforcement Learning System at Scale 18 Mar 2025 · 1 repository · arXiv:2503.14476
-
Free-Lunch Color-Texture Disentanglement for Stylized Image Generation 18 Mar 2025 · 0 repositories · arXiv:2503.14275
-
Organ-aware Multi-scale Medical Image Segmentation Using Text Prompt Engineering 18 Mar 2025 · 0 repositories · arXiv:2503.13806
-
SketchFusion: Learning Universal Sketch Features through Fusing Foundation Models 18 Mar 2025 · 0 repositories · arXiv:2503.14129
-
Evolution-based Region Adversarial Prompt Learning for Robustness Enhancement in Vision-Language Models 17 Mar 2025 · 0 repositories · arXiv:2503.12874
-
From Head to Tail: Towards Balanced Representation in Large Vision-Language Models through Adaptive Data Calibration 17 Mar 2025 · 0 repositories · arXiv:2503.12821
-
MFP-CLIP: Exploring the Efficacy of Multi-Form Prompts for Zero-Shot Industrial Anomaly Detection 17 Mar 2025 · 0 repositories · arXiv:2503.12910
-
Web Artifact Attacks Disrupt Vision Language Models 17 Mar 2025 · 0 repositories · arXiv:2503.13652
-
VRsketch2Gaussian: 3D VR Sketch Guided 3D Object Generation with Gaussian Splatting 16 Mar 2025 · 0 repositories · arXiv:2503.12383
-
Hyperbolic Safety-Aware Vision-Language Models 15 Mar 2025 · 1 repository · arXiv:2503.12127Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
LIAM: Multimodal Transformer for Language Instructions, Images, Actions and Semantic Maps 15 Mar 2025 · 1 repository · arXiv:2503.12230
-
Prosody-Enhanced Acoustic Pre-training and Acoustic-Disentangled Prosody Adapting for Movie Dubbing 15 Mar 2025 · 1 repository · arXiv:2503.12042
-
TLAC: Two-stage LMM Augmented CLIP for Zero-Shot Classification 15 Mar 2025 · 1 repository · arXiv:2503.12206
-
VTON 360: High-Fidelity Virtual Try-On from Any Viewing Direction 15 Mar 2025 · 0 repositories · arXiv:2503.12165
-
Quantifying Interpretability in CLIP Models with Concept Consistency 14 Mar 2025 · 0 repositories · arXiv:2503.11103
-
UStyle: Waterbody Style Transfer of Underwater Scenes by Depth-Guided Feature Synthesis 14 Mar 2025 · 1 repository · arXiv:2503.11893
-
4D LangSplat: 4D Language Gaussian Splatting via Multimodal Large Language Models 13 Mar 2025 · 1 repository · arXiv:2503.10437
-
A Hierarchical Semantic Distillation Framework for Open-Vocabulary Object Detection 13 Mar 2025 · 1 repository · arXiv:2503.10152
-
Technical Approach for the EMI Challenge in the 8th Affective Behavior Analysis in-the-Wild Competition 13 Mar 2025 · 0 repositories · arXiv:2503.10603
-
Emotion Recognition with CLIP and Sequential Learning 13 Mar 2025 · 0 repositories · arXiv:2503.09929
-
NeighborRetr: Balancing Hub Centrality in Cross-Modal Retrieval 13 Mar 2025 · 1 repository · arXiv:2503.10526
-
Team NYCU at Defactify4: Robust Detection and Source Identification of AI-Generated Images Using CNN and CLIP-Based Models 13 Mar 2025 · 1 repository · arXiv:2503.10718
-
Bayesian Test-Time Adaptation for Vision-Language Models 12 Mar 2025 · 0 repositories · arXiv:2503.09248
-
C^2 ATTACK: Towards Representation Backdoor on CLIP via Concept Confusion 12 Mar 2025 · 0 repositories · arXiv:2503.09095
-
Context-guided Responsible Data Augmentation with Diffusion Models 12 Mar 2025 · 1 repository · arXiv:2503.10687
-
On the Limitations of Vision-Language Models in Understanding Image Transforms 12 Mar 2025 · 0 repositories · arXiv:2503.09837
-
Online Language Splatting 12 Mar 2025 · 0 repositories · arXiv:2503.09447
-
Controlling Latent Diffusion Using Latent CLIP 11 Mar 2025 · 1 repository · arXiv:2503.08455
-
External Knowledge Injection for CLIP-Based Class-Incremental Learning 11 Mar 2025 · 3 repositories · arXiv:2503.08510
-
MMRL: Multi-Modal Representation Learning for Vision-Language Models 11 Mar 2025 · 1 repository · arXiv:2503.08497Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample)
-
Modeling Variants of Prompts for Vision-Language Models 11 Mar 2025 · 1 repository · arXiv:2503.08229
-
Prompt-OT: An Optimal Transport Regularization Paradigm for Knowledge Preservation in Vision-Language Model Adaptation 11 Mar 2025 · 1 repository · arXiv:2503.08906
-
CAPT: Class-Aware Prompt Tuning for Federated Long-Tailed Learning with Vision-Language Model 10 Mar 2025 · 0 repositories · arXiv:2503.06993
-
Is CLIP ideal? No. Can we fix it? Yes! 10 Mar 2025 · 1 repository · arXiv:2503.08723
-
Visual and Text Prompt Segmentation: A Novel Multi-Model Framework for Remote Sensing 10 Mar 2025 · 0 repositories · arXiv:2503.07911
-
WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation 10 Mar 2025 · 2 repositories · arXiv:2503.07265Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 4 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
AA-CLIP: Enhancing Zero-shot Anomaly Detection via Anomaly-Aware CLIP 9 Mar 2025 · 1 repository · arXiv:2503.06661
-
DiffCLIP: Differential Attention Meets CLIP 9 Mar 2025 · 1 repository · arXiv:2503.06626Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
M³amba: CLIP-driven Mamba Model for Multi-modal Remote Sensing Classification 9 Mar 2025 · 1 repository · arXiv:2503.06446
-
OT-DETECTOR: Delving into Optimal Transport for Zero-shot Out-of-Distribution Detection 9 Mar 2025 · 0 repositories · arXiv:2503.06442
-
SEED: Towards More Accurate Semantic Evaluation for Visual Brain Decoding 9 Mar 2025 · 0 repositories · arXiv:2503.06437Syntology 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples)
-
Integrating Frequency-Domain Representations with Low-Rank Adaptation in Vision-Language Models 8 Mar 2025 · 0 repositories · arXiv:2503.06003
-
Vision-aware Multimodal Prompt Tuning for Uploadable Multi-source Few-shot Domain Adaptation 8 Mar 2025 · 0 repositories · arXiv:2503.06106
-
Towards Locally Explaining Prediction Behavior via Gradual Interventions and Measuring Property Gradients 7 Mar 2025 · 0 repositories · arXiv:2503.05424
-
Inclusive STEAM Education: A Framework for Teaching Cod-2 ing and Robotics to Students with Visually Impairment Using 3 Advanced Computer Vision 6 Mar 2025 · 0 repositories · arXiv:2503.16482
-
CLIP is Strong Enough to Fight Back: Test-time Counterattacks towards Zero-shot Adversarial Robustness of CLIP 5 Mar 2025 · 1 repository · arXiv:2503.03613Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 2 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
UAR-NVC: A Unified AutoRegressive Framework for Memory-Efficient Neural Video Compression 4 Mar 2025 · 0 repositories · arXiv:2503.02733
-
Vision-Language Model IP Protection via Prompt-based Learning 4 Mar 2025 · 0 repositories · arXiv:2503.02393
-
ClipGrader: Leveraging Vision-Language Models for Robust Label Quality Assessment in Object Detection 3 Mar 2025 · 0 repositories · arXiv:2503.02897
-
Enhancing Vision-Language Compositional Understanding with Multimodal Synthetic Data 3 Mar 2025 · 0 repositories · arXiv:2503.01167
-
Generalizable Prompt Learning of CLIP: A Brief Overview 3 Mar 2025 · 0 repositories · arXiv:2503.01263
-
Language-Assisted Feature Transformation for Anomaly Detection 3 Mar 2025 · 1 repository · arXiv:2503.01184Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
OFF-CLIP: Improving Normal Detection Confidence in Radiology CLIP with Simple Off-Diagonal Term Auto-Adjustment 3 Mar 2025 · 1 repository · arXiv:2503.01794
-
Extrapolating and Decoupling Image-to-Video Generation Models: Motion Modeling is Easier Than You Think 2 Mar 2025 · 1 repository · arXiv:2503.00948Syntology official (archive's flag): 12 ran · 12 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 1 violated, 7 with no contract checked; 4 where Syntology's instrument failed) · 3 unverified (of 15 harvested samples) · 4 pointer-only (licence)
-
Quality-Driven Curation of Remote Sensing Vision-Language Data via Learned Scoring Models 2 Mar 2025 · 0 repositories · arXiv:2503.00743
-
Few-shot crack image classification using clip based on bayesian optimization 1 Mar 2025 · 0 repositories · arXiv:2503.00376
-
SGC-Net: Stratified Granular Comparison Network for Open-Vocabulary HOI Detection 1 Mar 2025 · 1 repository · arXiv:2503.00414
-
CoTMR: Chain-of-Thought Multi-Scale Reasoning for Training-Free Zero-Shot Composed Image Retrieval 28 Feb 2025 · 0 repositories · arXiv:2502.20826
-
UoR-NCL at SemEval-2025 Task 1: Using Generative LLMs and CLIP Models for Multilingual Multimodal Idiomaticity Representation 28 Feb 2025 · 0 repositories · arXiv:2502.20984
-
PET Image Denoising via Text-Guided Diffusion: Integrating Anatomical Priors through Text Prompts 28 Feb 2025 · 0 repositories · arXiv:2502.21260
-
Towards General Visual-Linguistic Face Forgery Detection(V2) 28 Feb 2025 · 1 repository · arXiv:2502.20698
-
Analyzing CLIP's Performance Limitations in Multi-Object Scenarios: A Controlled High-Resolution Study 27 Feb 2025 · 0 repositories · arXiv:2502.19828
-
Differential Contrastive Training for Gaze Estimation 27 Feb 2025 · 0 repositories · arXiv:2502.20128
-
CLIP Under the Microscope: A Fine-Grained Analysis of Multi-Object Representation 27 Feb 2025 · 1 repository · arXiv:2502.19842Syntology official (archive's flag): 4 ran · 5 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 2 honoured, 2 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
Interpreting CLIP with Hierarchical Sparse Autoencoders 27 Feb 2025 · 0 repositories · arXiv:2502.20578
-
Learning to Generalize without Bias for Open-Vocabulary Action Recognition 27 Feb 2025 · 0 repositories · arXiv:2502.20158Syntology 4 ran (of which 4 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified; every one of the 4 samples that ran constructed an object rather than computing a result (of 7 harvested samples)
-
Multimodal Representation Alignment for Image Generation: Text-Image Interleaved Control Is Easier Than You Think 27 Feb 2025 · 1 repository · arXiv:2502.20172
-
Open-Vocabulary Semantic Part Segmentation of 3D Human 27 Feb 2025 · 0 repositories · arXiv:2502.19782
-
UniTok: A Unified Tokenizer for Visual Generation and Understanding 27 Feb 2025 · 1 repository · arXiv:2502.20321Syntology official (archive's flag): 11 ran · 12 ran (of which 9 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 3 where Syntology's instrument failed) · 3 unverified (of 15 harvested samples) · 1 pointer-only (licence)
-
Attention-Guided Integration of CLIP and SAM for Precise Object Masking in Robotic Manipulation 26 Feb 2025 · 0 repositories · arXiv:2502.18842
-
Clip-TTS: Contrastive Text-content and Mel-spectrogram, A High-Quality Text-to-Speech Method based on Contextual Semantic Understanding 26 Feb 2025 · 0 repositories · arXiv:2502.18889
-
FAA-CLIP: Federated Adversarial Adaptation of CLIP 26 Feb 2025 · 1 repository · arXiv:2503.05776
-
CLIPure: Purification in Latent Space via CLIP for Adversarially Robust Zero-Shot Classification 25 Feb 2025 · 1 repository · arXiv:2502.18176Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
GCDance: Genre-Controlled 3D Full Body Dance Generation Driven By Music 25 Feb 2025 · 0 repositories · arXiv:2502.18309
-
LDGen: Enhancing Text-to-Image Synthesis via Large Language Model-Driven Language Representation 25 Feb 2025 · 0 repositories · arXiv:2502.18302
-
Zero-Shot Semantic Communication with Multimodal Foundation Models 25 Feb 2025 · 0 repositories · arXiv:2502.18200
-
CLIP-SENet: CLIP-based Semantic Enhancement Network for Vehicle Re-identification 24 Feb 2025 · 0 repositories · arXiv:2502.16815
-
Distributional Vision-Language Alignment by Cauchy-Schwarz Divergence 24 Feb 2025 · 0 repositories · arXiv:2502.17028
-
Category-Selective Neurons in Deep Networks: Comparing Purely Visual and Visual-Language Models 23 Feb 2025 · 0 repositories · arXiv:2502.16456
-
Dr. Splat: Directly Referring 3D Gaussian Splatting via Direct Language Embedding Registration 23 Feb 2025 · 0 repositories · arXiv:2502.16652
-
FeatSharp: Your Vision Model Features, Sharper 22 Feb 2025 · 1 repository · arXiv:2502.16025Syntology 11 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 2 where Syntology's instrument failed) · 9 unverified (of 20 harvested samples) · 20 pointer-only (licence)
-
ELIP: Enhanced Visual-Language Foundation Models for Image Retrieval 21 Feb 2025 · 0 repositories · arXiv:2502.15682
-
Modality-Aware Neuron Pruning for Unlearning in Multimodal Large Language Models 21 Feb 2025 · 1 repository · arXiv:2502.15910Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 6 where Syntology's instrument failed) · 12 unverified (of 18 harvested samples) · 18 pointer-only (licence)
-
TransMamba: Fast Universal Architecture Adaption from Transformers to Mamba 21 Feb 2025 · 0 repositories · arXiv:2502.15130
-
Visual Zero-Shot E-Commerce Product Attribute Value Extraction 21 Feb 2025 · 0 repositories · arXiv:2502.15979
-
Hardware-Friendly Static Quantization Method for Video Diffusion Transformers 20 Feb 2025 · 0 repositories · arXiv:2502.15077
-
Enhancing Chest X-ray Classification through Knowledge Injection in Cross-Modality Learning 19 Feb 2025 · 0 repositories · arXiv:2502.13447
-
Generative Video Semantic Communication via Multimodal Semantic Fusion with Large Model 19 Feb 2025 · 0 repositories · arXiv:2502.13838
-
IP-Composer: Semantic Composition of Visual Concepts 19 Feb 2025 · 0 repositories · arXiv:2502.13951
-
Object-centric Binding in Contrastive Language-Image Pretraining 19 Feb 2025 · 0 repositories · arXiv:2502.14113
-
AttriVision: Advancing Generalization in Pedestrian Attribute Recognition using CLIP 18 Feb 2025 · 0 repositories
-
RealSyn: An Effective and Scalable Multimodal Interleaved Document Transformation Paradigm 18 Feb 2025 · 1 repository · arXiv:2502.12513
-
SongGen: A Single Stage Auto-regressive Transformer for Text-to-Song Generation 18 Feb 2025 · 1 repository · arXiv:2502.13128Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 1 honoured, 1 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Adversarially Robust CLIP Models Can Induce Better (Robust) Perceptual Metrics 17 Feb 2025 · 1 repository · arXiv:2502.11725
-
Control-CLIP: Decoupling Category and Style Guidance in CLIP for Specific-Domain Generation 17 Feb 2025 · 0 repositories · arXiv:2502.11532
-
Descriminative-Generative Custom Tokens for Vision-Language Models 17 Feb 2025 · 0 repositories · arXiv:2502.12095
-
GeoDANO: Geometric VLM with Domain Agnostic Vision Encoder 17 Feb 2025 · 0 repositories · arXiv:2502.11360
-
Pretraining Frequency Predicts Compositional Generalization of CLIP on Real-World Tasks 17 Feb 2025 · 0 repositories · arXiv:2502.18326