Methods › Computer Vision › Vision and Language Pre-Trained Models › CLIP › Papers, page 9
Contrastive Language-Image Pre-training
CLIP
Papers archive 2025-07-28
archive papers tagged: 3,094 · with a code link: 1,617 · where Syntology ran a sample: 649 (554 with a run with no instrument failure, 95 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (649 of 3,094 tagged: 554 with a run with no instrument failure, 95 where every run was a failure of Syntology's instrument)
Page 9 of 31: papers 801 to 900 of 3,094, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
A Hybrid Defense Strategy for Boosting Adversarial Robustness in Vision-Language Models 18 Oct 2024 · 0 repositories · arXiv:2410.14911
-
CLIP-VAD: Exploiting Vision-Language Models for Voice Activity Detection 18 Oct 2024 · 0 repositories · arXiv:2410.14509
-
LUDVIG: Learning-free Uplifting of 2D Visual features to Gaussian Splatting scenes 18 Oct 2024 · 0 repositories · arXiv:2410.14462
-
NaturalBench: Evaluating Vision-Language Models on Natural Adversarial Samples 18 Oct 2024 · 0 repositories · arXiv:2410.14669
-
ConsisSR: Delving Deep into Consistency in Diffusion-based Image Super-Resolution 17 Oct 2024 · 0 repositories · arXiv:2410.13807
-
Improving Multi-modal Large Language Model through Boosting Vision Capabilities 17 Oct 2024 · 0 repositories · arXiv:2410.13733
-
Learning Multimodal Cues of Children's Uncertainty 17 Oct 2024 · 0 repositories · arXiv:2410.14050
-
Let Me Finish My Sentence: Video Temporal Grounding with Holistic Text Understanding 17 Oct 2024 · 0 repositories · arXiv:2410.13598
-
debiaSAE: Benchmarking and Mitigating Vision-Language Model Bias 17 Oct 2024 · 1 repository · arXiv:2410.13146
-
Performance of Gaussian Mixture Model Classifiers on Embedded Feature Spaces 17 Oct 2024 · 1 repository · arXiv:2410.13421
-
DH-VTON: Deep Text-Driven Virtual Try-On via Hybrid Attention Learning 16 Oct 2024 · 0 repositories · arXiv:2410.12501
-
FTII-Bench: A Comprehensive Multimodal Benchmark for Flow Text with Image Insertion 16 Oct 2024 · 1 repository · arXiv:2410.12564
-
Hiding-in-Plain-Sight (HiPS) Attack on CLIP for Targetted Object Removal from Images 16 Oct 2024 · 0 repositories · arXiv:2410.13010
-
Interpreting and Analysing CLIP's Zero-Shot Image Classification via Mutual Knowledge 16 Oct 2024 · 1 repository · arXiv:2410.13016Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Mind the Gap Between Prototypes and Images in Cross-domain Finetuning 16 Oct 2024 · 1 repository · arXiv:2410.12474
-
Is Less More? Exploring Token Condensation as Training-free Adaptation for CLIP 16 Oct 2024 · 1 repository · arXiv:2410.14729
-
TransAgent: Transfer Vision-Language Foundation Models with Heterogeneous Agent Collaboration 16 Oct 2024 · 1 repository · arXiv:2410.12183Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 3 where Syntology's instrument failed) · 2 unverified (of 8 harvested samples) · 2 pointer-only (licence)
-
Preserve or Modify? Context-Aware Evaluation for Balancing Preservation and Modification in Text-Guided Image Editing 15 Oct 2024 · 0 repositories · arXiv:2410.11374
-
CLIP-DFGS: A Hard Sample Mining Method for CLIP in Generalizable Person Re-Identification 15 Oct 2024 · 0 repositories · arXiv:2410.11255
-
CtrlSynth: Controllable Image Text Synthesis for Data-Efficient Multimodal Learning 15 Oct 2024 · 0 repositories · arXiv:2410.11963
-
Improving Long-Text Alignment for Text-to-Image Diffusion Models 15 Oct 2024 · 1 repository · arXiv:2410.11817Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 1 where Syntology's instrument failed) · 6 unverified (of 10 harvested samples) · 1 pointer-only (licence)
-
It's Just Another Day: Unique Video Captioning by Discriminative Prompting 15 Oct 2024 · 0 repositories · arXiv:2410.11702
-
Animate-X: Universal Character Image Animation with Enhanced Motion Representation 14 Oct 2024 · 0 repositories · arXiv:2410.10306
-
HART: Efficient Visual Generation with Hybrid Autoregressive Transformer 14 Oct 2024 · 2 repositories · arXiv:2410.10812Syntology official (archive's flag): 23 ran · 23 ran (of which 0 constructed an object rather than computing a result; 15 with no instrument failure: 2 honoured, 0 violated, 13 with no contract checked; 8 where Syntology's instrument failed) · 4 unverified (of 27 harvested samples) · 4 pointer-only (licence)
-
Learning to Customize Text-to-Image Diffusion In Diverse Context 14 Oct 2024 · 0 repositories · arXiv:2410.10058
-
LOBG:Less Overfitting for Better Generalization in Vision-Language Model 14 Oct 2024 · 0 repositories · arXiv:2410.10247
-
Locality Alignment Improves Vision-Language Models 14 Oct 2024 · 1 repository · arXiv:2410.11087Syntology 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Mixture of Experts Made Personalized: Federated Prompt Learning for Vision-Language Models 14 Oct 2024 · 1 repository · arXiv:2410.10114Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples) · 2 pointer-only (licence)
-
MMAR: Towards Lossless Multi-Modal Auto-Regressive Probabilistic Modeling 14 Oct 2024 · 0 repositories · arXiv:2410.10798
-
Generating Intermediate Representations for Compositional Text-To-Image Generation 13 Oct 2024 · 1 repository · arXiv:2410.09792Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
TULIP: Token-length Upgraded CLIP 13 Oct 2024 · 1 repository · arXiv:2410.10034Syntology official (archive's flag): 13 ran · 13 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 1 violated, 8 with no contract checked; 4 where Syntology's instrument failed) · 7 unverified (of 20 harvested samples) · 7 pointer-only (licence)
-
CLIP-SCGI: Synthesized Caption-Guided Inversion for Person Re-Identification 12 Oct 2024 · 0 repositories · arXiv:2410.09382
-
Debiasing Vison-Language Models with Text-Only Training 12 Oct 2024 · 0 repositories · arXiv:2410.09365
-
Boosting Open-Vocabulary Object Detection by Handling Background Samples 11 Oct 2024 · 0 repositories · arXiv:2410.08645
-
Calibrated Cache Model for Few-Shot Vision-Language Model Adaptation 11 Oct 2024 · 0 repositories · arXiv:2410.08895
-
More than Memes: A Multimodal Topic Modeling Approach to Conspiracy Theories on Telegram 11 Oct 2024 · 0 repositories · arXiv:2410.08642
-
Natural Language Induced Adversarial Images 11 Oct 2024 · 1 repository · arXiv:2410.08620
-
Semantic Token Reweighting for Interpretable and Controllable Text Embeddings in CLIP 11 Oct 2024 · 0 repositories · arXiv:2410.08469
-
CLIP Multi-modal Hashing for Multimedia Retrieval 10 Oct 2024 · 0 repositories · arXiv:2410.07783
-
CSA: Data-efficient Mapping of Unimodal Features to Multimodal Features 10 Oct 2024 · 0 repositories · arXiv:2410.07610
-
Emerging Pixel Grounding in Large Multimodal Models Without Grounding Supervision 10 Oct 2024 · 0 repositories · arXiv:2410.08209
-
FLIER: Few-shot Language Image Models Embedded with Latent Representations 10 Oct 2024 · 0 repositories · arXiv:2410.07648
-
HeGraphAdapter: Tuning Multi-Modal Vision-Language Models with Heterogeneous Graph Adapter 10 Oct 2024 · 0 repositories · arXiv:2410.07854
-
In Search of Forgotten Domain Generalization 10 Oct 2024 · 0 repositories · arXiv:2410.08258
-
LatteCLIP: Unsupervised CLIP Fine-Tuning via LMM-Synthetic Texts 10 Oct 2024 · 0 repositories · arXiv:2410.08211
-
Progressive Autoregressive Video Diffusion Models 10 Oct 2024 · 1 repository · arXiv:2410.08151Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 7 harvested samples) · 2 pointer-only (licence)
-
Boosting Few-Shot Detection with Large Language Models and Layout-to-Image Synthesis 9 Oct 2024 · 0 repositories · arXiv:2410.06841
-
Compositional Entailment Learning for Hyperbolic Vision-Language Models 9 Oct 2024 · 1 repository · arXiv:2410.06912Syntology official (archive's flag): 7 ran · 7 ran (of which 2 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 5 where Syntology's instrument failed) · 6 unverified (of 13 harvested samples) · 13 pointer-only (licence)
-
Enhancing Vision-Language Model Pre-training with Image-text Pair Pruning Based on Word Frequency 9 Oct 2024 · 1 repository · arXiv:2410.10879
-
Exploiting Distribution Constraints for Scalable and Efficient Image Retrieval 9 Oct 2024 · 0 repositories · arXiv:2410.07022
-
Fostering Intrinsic Motivation in Reinforcement Learning with Pretrained Foundation Models 9 Oct 2024 · 0 repositories · arXiv:2410.07404
-
LaMP: Language-Motion Pretraining for Motion Generation, Retrieval, and Captioning 9 Oct 2024 · 0 repositories · arXiv:2410.07093
-
Open-RGBT: Open-vocabulary RGB-T Zero-shot Semantic Segmentation in Open-world Environments 9 Oct 2024 · 0 repositories · arXiv:2410.06626
-
Positive-Augmented Contrastive Learning for Vision-and-Language Evaluation and Training 9 Oct 2024 · 1 repository · arXiv:2410.07336
-
CASA: Class-Agnostic Shared Attributes in Vision-Language Models for Efficient Incremental Object Detection 8 Oct 2024 · 0 repositories · arXiv:2410.05804
-
FACMIC: Federated Adaptative CLIP Model for Medical Image Classification 8 Oct 2024 · 1 repository · arXiv:2410.14707
-
PixLens: A Novel Framework for Disentangled Evaluation in Diffusion-Based Image Editing with Object Detection + SAM 8 Oct 2024 · 1 repository · arXiv:2410.05710
-
SIA-OVD: Shape-Invariant Adapter for Bridging the Image-Region Gap in Open-Vocabulary Detection 8 Oct 2024 · 1 repository · arXiv:2410.05650
-
Fine-Tuning CLIP's Last Visual Projector: A Few-Shot Cornucopia 7 Oct 2024 · 1 repository · arXiv:2410.05270
-
Preserving Multi-Modal Capabilities of Pre-trained VLMs for Improving Vision-Linguistic Compositionality 7 Oct 2024 · 1 repository · arXiv:2410.05210Syntology official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 1 honoured, 1 violated, 1 with no contract checked; 7 where Syntology's instrument failed) · 3 unverified (of 13 harvested samples) · 13 pointer-only (licence)
-
SELECT: A Large-Scale Benchmark of Data Curation Strategies for Image Classification 7 Oct 2024 · 1 repository · arXiv:2410.05057Syntology official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 13 harvested samples)
-
VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding Tasks 7 Oct 2024 · 0 repositories · arXiv:2410.05160
-
Zero-Shot Vision-and-Language Navigation with Collision Mitigation in Continuous Environment 7 Oct 2024 · 0 repositories · arXiv:2410.17267
-
Multimodal 3D Fusion and In-Situ Learning for Spatially Aware AI 6 Oct 2024 · 1 repository · arXiv:2410.04652
-
Learning on LoRAs: GL-Equivariant Processing of Low-Rank Weight Spaces for Large Finetuned Models 5 Oct 2024 · 0 repositories · arXiv:2410.04207
-
Combing Text-based and Drag-based Editing for Precise and Flexible Image Editing 4 Oct 2024 · 0 repositories · arXiv:2410.03097
-
DiSK: Differentially Private Optimizer with Simplified Kalman Filter for Noise Reduction 4 Oct 2024 · 0 repositories · arXiv:2410.03883
-
Generalizable Prompt Tuning for Vision-Language Models 4 Oct 2024 · 0 repositories · arXiv:2410.03189
-
Investigating and Mitigating Object Hallucinations in Pretrained Vision-Language (CLIP) Models 4 Oct 2024 · 1 repository · arXiv:2410.03176Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
A Retention-Centric Framework for Continual Learning with Guaranteed Model Developmental Safety 4 Oct 2024 · 1 repository · arXiv:2410.03955Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 4 where Syntology's instrument failed) · 4 unverified (of 9 harvested samples) · 9 pointer-only (licence)
-
Redefining Temporal Modeling in Video Diffusion: The Vectorized Timestep Approach 4 Oct 2024 · 1 repository · arXiv:2410.03160Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 1 honoured, 0 violated, 6 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 10 harvested samples)
-
ShieldDiff: Suppressing Sexual Content Generation from Diffusion Models through Reinforcement Learning 4 Oct 2024 · 0 repositories · arXiv:2410.05309
-
The Wallpaper is Ugly: Indoor Localization using Vision and Language 4 Oct 2024 · 0 repositories · arXiv:2410.03900
-
VEDIT: Latent Prediction Architecture For Procedural Video Representation Learning 4 Oct 2024 · 0 repositories · arXiv:2410.03478
-
Contrastive Localized Language-Image Pre-Training 3 Oct 2024 · 0 repositories · arXiv:2410.02746
-
Revisit Large-Scale Image-Caption Data in Pre-training Multimodal Foundation Models 3 Oct 2024 · 1 repository · arXiv:2410.02740
-
Understanding and Mitigating Miscalibration in Prompt Tuning for Vision-Language Models 3 Oct 2024 · 1 repository · arXiv:2410.02681Syntology 4 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 3 where Syntology's instrument failed) · 5 unverified (of 9 harvested samples) · 6 pointer-only (licence)
-
AgriCLIP: Adapting CLIP for Agriculture and Livestock via Domain-Specialized Cross-Model Alignment 2 Oct 2024 · 1 repository · arXiv:2410.01407
-
Edge-preserving noise for diffusion models 2 Oct 2024 · 1 repository · arXiv:2410.01540Syntology 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Integrating Visual and Textual Inputs for Searching Large-Scale Map Collections with CLIP 2 Oct 2024 · 1 repository · arXiv:2410.01190
-
SegEarth-OV: Towards Training-Free Open-Vocabulary Segmentation for Remote Sensing Images 2 Oct 2024 · 2 repositories · arXiv:2410.01768Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 1 violated, 3 with no contract checked; 2 where Syntology's instrument failed) · 3 unverified (of 9 harvested samples) · 9 pointer-only (licence)
-
Toward a Holistic Evaluation of Robustness in CLIP Models 2 Oct 2024 · 0 repositories · arXiv:2410.01534
-
PointAD: Comprehending 3D Anomalies from Points and Pixels for Zero-shot 3D Anomaly Detection 1 Oct 2024 · 1 repository · arXiv:2410.00320Syntology official (archive's flag): 15 ran · 17 ran (of which 2 constructed an object rather than computing a result; 11 with no instrument failure: 1 honoured, 0 violated, 10 with no contract checked; 6 where Syntology's instrument failed) · 9 unverified (of 26 harvested samples) · 23 pointer-only (licence)
-
Rethinking Misalignment in Vision-Language Model Adaptation from a Causal Perspective 1 Oct 2024 · 0 repositories · arXiv:2410.12816
-
VideoCLIP-XL: Advancing Long Description Understanding for Video CLIP Models 1 Oct 2024 · 0 repositories · arXiv:2410.00741
-
Magnet: We Never Know How Text-to-Image Diffusion Models Work, Until We Learn How Vision-Language Models Function 30 Sep 2024 · 1 repository · arXiv:2409.19967Syntology official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 1 honoured, 0 violated, 9 with no contract checked; 1 where Syntology's instrument failed) · 6 unverified (of 17 harvested samples) · 17 pointer-only (licence)
-
ProFD: Prompt-Guided Feature Disentangling for Occluded Person Re-Identification 30 Sep 2024 · 1 repository · arXiv:2409.20081
-
The age of spiritual machines: Language quietus induces synthetic altered states of consciousness in artificial intelligence 30 Sep 2024 · 0 repositories · arXiv:2410.00257
-
Towards Open-Vocabulary Semantic Segmentation Without Semantic Labels 30 Sep 2024 · 0 repositories · arXiv:2409.19846
-
TROPE: TRaining-Free Object-Part Enhancement for Seamlessly Improving Fine-Grained Zero-Shot Image Captioning 30 Sep 2024 · 1 repository · arXiv:2409.19960
-
CLIP-based Camera-Agnostic Feature Learning for Intra-camera Person Re-Identification 29 Sep 2024 · 1 repository · arXiv:2409.19563
-
Efficient Backdoor Defense in Multimodal Contrastive Learning: A Token-Level Unlearning Method for Mitigating Threats 29 Sep 2024 · 0 repositories · arXiv:2409.19526
-
Federated Learning from Vision-Language Foundation Models: Theoretical Analysis and Method 29 Sep 2024 · 1 repository · arXiv:2409.19610Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 4 where Syntology's instrument failed) · 3 unverified (of 10 harvested samples) · 10 pointer-only (licence)
-
CLIP-MoE: Towards Building Mixture of Experts for CLIP with Diversified Multiplet Upcycling 28 Sep 2024 · 1 repository · arXiv:2409.19291Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
DOTA: Distributional Test-Time Adaptation of Vision-Language Models 28 Sep 2024 · 0 repositories · arXiv:2409.19375
-
FairPIVARA: Reducing and Assessing Biases in CLIP-Based Multimodal Models 28 Sep 2024 · 1 repository · arXiv:2409.19474
-
From Unimodal to Multimodal: Scaling up Projectors to Align Modalities 28 Sep 2024 · 1 repository · arXiv:2409.19425
-
MedCLIP-SAMv2: Towards Universal Text-Driven Medical Image Segmentation 28 Sep 2024 · 2 repositories · arXiv:2409.19483
-
Emu3: Next-Token Prediction is All You Need 27 Sep 2024 · 2 repositories · arXiv:2409.18869Syntology 6 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 6 harvested samples)
-
Improving Visual Object Tracking through Visual Prompting 27 Sep 2024 · 1 repository · arXiv:2409.18901