Methods › Computer Vision › Vision and Language Pre-Trained Models › CLIP › Papers, page 10
Contrastive Language-Image Pre-training
CLIP
Papers archive 2025-07-28
archive papers tagged: 3,094 · with a code link: 1,617 · where Syntology ran a sample: 649 (554 with a run with no instrument failure, 95 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (649 of 3,094 tagged: 554 with a run with no instrument failure, 95 where every run was a failure of Syntology's instrument)
Page 10 of 31: papers 901 to 1,000 of 3,094, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Individuation in Neural Models with and without Visual Grounding 27 Sep 2024 · 0 repositories · arXiv:2409.18868
-
UniEmoX: Cross-modal Semantic-Guided Large-Scale Pretraining for Universal Scene Emotion Perception 27 Sep 2024 · 1 repository · arXiv:2409.18877
-
Cascade Prompt Learning for Vision-Language Model Adaptation 26 Sep 2024 · 2 repositories · arXiv:2409.17805
-
Wavelet-Driven Generalizable Framework for Deepfake Face Forgery Detection 26 Sep 2024 · 1 repository · arXiv:2409.18301
-
LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness 26 Sep 2024 · 0 repositories · arXiv:2409.18125
-
Multi-View and Multi-Scale Alignment for Contrastive Language-Image Pre-training in Mammography 26 Sep 2024 · 1 repository · arXiv:2409.18119
-
MultiClimate: Multimodal Stance Detection on Climate Change Videos 26 Sep 2024 · 1 repository · arXiv:2409.18346Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Pioneering Reliable Assessment in Text-to-Image Knowledge Editing: Leveraging a Fine-Grained Dataset and an Innovative Criterion 26 Sep 2024 · 1 repository · arXiv:2409.17928
-
Robotic-CLIP: Fine-tuning CLIP on Action Data for Robotic Applications 26 Sep 2024 · 0 repositories · arXiv:2409.17727
-
CleanerCLIP: Fine-grained Counterfactual Semantic Augmentation for Backdoor Defense in Contrastive Learning 26 Sep 2024 · 0 repositories · arXiv:2409.17601
-
The Hard Positive Truth about Vision-Language Compositionality 26 Sep 2024 · 1 repository · arXiv:2409.17958Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Attention Prompting on Image for Large Vision-Language Models 25 Sep 2024 · 1 repository · arXiv:2409.17143Syntology official (archive's flag): 6 ran · 8 ran (of which 3 constructed an object rather than computing a result; 5 with no instrument failure: 2 honoured, 0 violated, 3 with no contract checked; 3 where Syntology's instrument failed) · 4 unverified (of 12 harvested samples) · 2 pointer-only (licence)
-
Vision-Language Model Fine-Tuning via Simple Parameter-Efficient Modification 25 Sep 2024 · 1 repository · arXiv:2409.16718Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 3 where Syntology's instrument failed) · 2 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
Adversarial Backdoor Defense in CLIP 24 Sep 2024 · 0 repositories · arXiv:2409.15968
-
Lessons and Insights from a Unifying Study of Parameter-Efficient Fine-Tuning (PEFT) in Visual Recognition 24 Sep 2024 · 2 repositories · arXiv:2409.16434
-
MIMO: Controllable Character Video Synthesis with Spatial Decomposed Modeling 24 Sep 2024 · 0 repositories · arXiv:2409.16160
-
Can CLIP Count Stars? An Empirical Study on Quantity Bias in CLIP 23 Sep 2024 · 0 repositories · arXiv:2409.15035
-
Exploring Fine-grained Retail Product Discrimination with Zero-shot Object Classification Using Vision-Language Models 23 Sep 2024 · 0 repositories · arXiv:2409.14963
-
MemeCLIP: Leveraging CLIP Representations for Multimodal Meme Classification 23 Sep 2024 · 1 repository · arXiv:2409.14703
-
MIMAFace: Face Animation via Motion-Identity Modulated Appearance Feature Learning 23 Sep 2024 · 0 repositories · arXiv:2409.15179
-
TSCLIP: Robust CLIP Fine-Tuning for Worldwide Cross-Regional Traffic Sign Recognition 23 Sep 2024 · 1 repository · arXiv:2409.15077
-
VLEU: a Method for Automatic Evaluation for Generalizability of Text-to-Image Models 23 Sep 2024 · 1 repository · arXiv:2409.14704Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample)
-
Patch Ranking: Efficient CLIP by Learning to Rank Local Patches 22 Sep 2024 · 1 repository · arXiv:2409.14607
-
Self-Supervised Audio-Visual Soundscape Stylization 22 Sep 2024 · 0 repositories · arXiv:2409.14340
-
CUS3D :CLIP-based Unsupervised 3D Segmentation via Object-level Denoise 21 Sep 2024 · 0 repositories · arXiv:2409.13982
-
PromptTA: Prompt-driven Text Adapter for Source-free Domain Generalization 21 Sep 2024 · 1 repository · arXiv:2409.14163
-
DAP-LED: Learning Degradation-Aware Priors with CLIP for Joint Low-light Enhancement and Deblurring 20 Sep 2024 · 0 repositories · arXiv:2409.13496
-
Embedding Geometries of Contrastive Language-Image Pre-Training 19 Sep 2024 · 1 repository · arXiv:2409.13079Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 1 violated, 5 with no contract checked; 3 where Syntology's instrument failed) · 6 unverified (of 15 harvested samples) · 15 pointer-only (licence)
-
End-to-end Open-vocabulary Video Visual Relationship Detection using Multi-modal Prompting 19 Sep 2024 · 0 repositories · arXiv:2409.12499
-
ABHINAW: A method for Automatic Evaluation of Typography within AI-Generated Images 18 Sep 2024 · 0 repositories · arXiv:2409.11874
-
Designing Interfaces for Multimodal Vector Search Applications 18 Sep 2024 · 0 repositories · arXiv:2409.11629
-
Knowledge Adaptation Network for Few-Shot Class-Incremental Learning 18 Sep 2024 · 0 repositories · arXiv:2409.11770
-
Mixture of Prompt Learning for Vision Language Models 18 Sep 2024 · 0 repositories · arXiv:2409.12011
-
Pareto Data Framework: Steps Towards Resource-Efficient Decision Making Using Minimum Viable Data (MVD) 18 Sep 2024 · 0 repositories · arXiv:2409.12112
-
CLIP Adaptation by Intra-modal Overlap Reduction 17 Sep 2024 · 0 repositories · arXiv:2409.11338
-
Improving the Efficiency of Visually Augmented Language Models 17 Sep 2024 · 1 repository · arXiv:2409.11148
-
Less is More: A Simple yet Effective Token Reduction Method for Efficient Multi-modal LLMs 17 Sep 2024 · 1 repository · arXiv:2409.10994Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 1 honoured, 1 violated, 2 with no contract checked; 4 where Syntology's instrument failed) · 0 unverified (of 8 harvested samples) · 1 pointer-only (licence)
-
Multimodal Attention-Enhanced Feature Fusion-based Weekly Supervised Anomaly Violence Detection 17 Sep 2024 · 0 repositories · arXiv:2409.11223
-
Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models 16 Sep 2024 · 1 repository · arXiv:2409.10695Syntology 18 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 4 honoured, 1 violated, 4 with no contract checked; 9 where Syntology's instrument failed) · 3 unverified (of 21 harvested samples) · 21 pointer-only (licence)
-
Bias Begets Bias: The Impact of Biased Embeddings on Diffusion Models 15 Sep 2024 · 0 repositories · arXiv:2409.09569
-
Can Large Language Models Grasp Event Signals? Exploring Pure Zero-Shot Event-based Recognition 15 Sep 2024 · 1 repository · arXiv:2409.09628
-
Finetuning CLIP to Reason about Pairwise Differences 15 Sep 2024 · 1 repository · arXiv:2409.09721
-
MFCLIP: Multi-modal Fine-grained CLIP for Generalizable Diffusion Face Forgery Detection 15 Sep 2024 · 1 repository · arXiv:2409.09724
-
Detect Fake with Fake: Leveraging Synthetic Data-driven Representation for Synthetic Image Detection 13 Sep 2024 · 1 repository · arXiv:2409.08884
-
ComAlign: Compositional Alignment in Vision-Language Models 12 Sep 2024 · 0 repositories · arXiv:2409.08206
-
Dynamic Prompting of Frozen Text-to-Image Diffusion Models for Panoptic Narrative Grounding 12 Sep 2024 · 0 repositories · arXiv:2409.08251
-
Improving Virtual Try-On with Garment-focused Diffusion Models 12 Sep 2024 · 1 repository · arXiv:2409.08258Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 1 honoured, 0 violated, 5 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 8 harvested samples) · 8 pointer-only (licence)
-
Rethinking Prompting Strategies for Multi-Label Recognition with Partial Annotations 12 Sep 2024 · 0 repositories · arXiv:2409.08381
-
Top-down Activity Representation Learning for Video Question Answering 12 Sep 2024 · 0 repositories · arXiv:2409.07748
-
Multimodal Emotion Recognition with Vision-language Prompting and Modality Dropout 11 Sep 2024 · 0 repositories · arXiv:2409.07078
-
Realistic and Efficient Face Swapping: A Unified Approach with Diffusion Models 11 Sep 2024 · 1 repository · arXiv:2409.07269
-
Securing Vision-Language Models with a Robust Encoder Against Jailbreak and Adversarial Attacks 11 Sep 2024 · 0 repositories · arXiv:2409.07353
-
DACAT: Dual-stream Adaptive Clip-aware Time Modeling for Robust Online Surgical Phase Recognition 10 Sep 2024 · 1 repository · arXiv:2409.06217
-
DetailCLIP: Detail-Oriented CLIP for Fine-Grained Tasks 10 Sep 2024 · 1 repository · arXiv:2409.06809Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 4 where Syntology's instrument failed) · 5 unverified (of 11 harvested samples) · 11 pointer-only (licence)
-
DiffQRCoder: Diffusion-based Aesthetic QR Code Generation with Scanning Robustness Guided Iterative Refinement 10 Sep 2024 · 1 repository · arXiv:2409.06355
-
ExIQA: Explainable Image Quality Assessment Using Distortion Attributes 10 Sep 2024 · 0 repositories · arXiv:2409.06853
-
Quantifying and Enabling the Interpretability of CLIP-like Models 10 Sep 2024 · 0 repositories · arXiv:2409.06579
-
Boosting CLIP Adaptation for Image Quality Assessment via Meta-Prompt Learning and Gradient Regularization 9 Sep 2024 · 0 repositories · arXiv:2409.05381
-
BrainDecoder: Style-Based Visual Decoding of EEG Signals 9 Sep 2024 · 0 repositories · arXiv:2409.05279
-
TriplePlay: Enhancing Federated Learning with CLIP for Non-IID Data and Resource Efficiency 9 Sep 2024 · 0 repositories · arXiv:2409.05347
-
FrozenSeg: Harmonizing Frozen Foundation Models for Open-Vocabulary Segmentation 5 Sep 2024 · 0 repositories · arXiv:2409.03525
-
Have Large Vision-Language Models Mastered Art History? 5 Sep 2024 · 0 repositories · arXiv:2409.03521
-
Text-Guided Mixup Towards Long-Tailed Image Categorization 5 Sep 2024 · 1 repository · arXiv:2409.03583
-
Standing on the Shoulders of Giants: Reprogramming Visual-Language Model for General Deepfake Detection 4 Sep 2024 · 0 repositories · arXiv:2409.02664
-
Evaluation and Comparison of Visual Language Models for Transportation Engineering Problems 3 Sep 2024 · 1 repository · arXiv:2409.02278
-
Multi-Modal Adapter for Vision-Language Models 3 Sep 2024 · 1 repository · arXiv:2409.02958
-
Optimizing CLIP Models for Image Retrieval with Maintained Joint-Embedding Alignment 3 Sep 2024 · 1 repository · arXiv:2409.01936Syntology official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 11 with no instrument failure: 0 honoured, 3 violated, 8 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 14 harvested samples) · 11 pointer-only (licence)
-
Taming CLIP for Fine-grained and Structured Visual Understanding of Museum Exhibits 3 Sep 2024 · 1 repository · arXiv:2409.01690
-
TempMe: Video Temporal Token Merging for Efficient Text-Video Retrieval 2 Sep 2024 · 1 repository · arXiv:2409.01156Syntology 7 ran (of which 3 constructed an object rather than computing a result; 4 with no instrument failure: 1 honoured, 0 violated, 3 with no contract checked; 3 where Syntology's instrument failed) · 1 unverified (of 8 harvested samples)
-
YOLOO: You Only Learn from Others Once 1 Sep 2024 · 0 repositories · arXiv:2409.00618
-
Aligning Medical Images with General Knowledge from Large Language Models 31 Aug 2024 · 1 repository · arXiv:2409.00341
-
COSMo: CLIP Talks on Open-Set Multi-Target Domain Adaptation 31 Aug 2024 · 1 repository · arXiv:2409.00397
-
EraseDraw: Learning to Draw Step-by-Step via Erasing Objects from Images 31 Aug 2024 · 0 repositories · arXiv:2409.00522
-
FADE: Few-shot/zero-shot Anomaly Detection Engine using Large Vision-Language Model 31 Aug 2024 · 1 repository · arXiv:2409.00556
-
TSO: Self-Training with Scaled Preference Optimization 31 Aug 2024 · 0 repositories · arXiv:2409.02118
-
AWRaCLe: All-Weather Image Restoration using Visual In-Context Learning 30 Aug 2024 · 0 repositories · arXiv:2409.00263
-
Retrieval-Augmented Natural Language Reasoning for Explainable Visual Question Answering 30 Aug 2024 · 0 repositories · arXiv:2408.17006
-
Text-to-Image Generation Via Energy-Based CLIP 30 Aug 2024 · 0 repositories · arXiv:2408.17046
-
Enhancing Conditional Image Generation with Explainable Latent Space Manipulation 29 Aug 2024 · 1 repository · arXiv:2408.16232
-
Fluent and Accurate Image Captioning with a Self-Trained Reward Model 29 Aug 2024 · 0 repositories · arXiv:2408.16827
-
A Simple Baseline with Single-encoder for Referring Image Segmentation 28 Aug 2024 · 0 repositories · arXiv:2408.15521
-
CoRe: Context-Regularized Text Embedding Learning for Text-to-Image Personalization 28 Aug 2024 · 0 repositories · arXiv:2408.15914
-
DEAR: Depth-Enhanced Action Recognition 28 Aug 2024 · 1 repository · arXiv:2408.15679
-
DiffAge3D: Diffusion-based 3D-aware Face Aging 28 Aug 2024 · 0 repositories · arXiv:2408.15922
-
More Text, Less Point: Towards 3D Data-Efficient Point-Language Understanding 28 Aug 2024 · 1 repository · arXiv:2408.15966
-
Perceive-IR: Learning to Perceive Degradation Better for All-in-One Image Restoration 28 Aug 2024 · 1 repository · arXiv:2408.15994
-
Visual Prompt Engineering for Medical Vision Language Models in Radiology 28 Aug 2024 · 0 repositories · arXiv:2408.15802
-
CLIP-AGIQA: Boosting the Performance of AI-Generated Image Quality Assessment with CLIP 27 Aug 2024 · 0 repositories · arXiv:2408.15098
-
From Bias to Balance: Detecting Facial Expression Recognition Biases in Large Multimodal Foundation Models 27 Aug 2024 · 0 repositories · arXiv:2408.14842
-
HPT++: Hierarchically Prompting Vision-Language Models with Multi-Granularity Knowledge Generation and Improved Structure Modeling 27 Aug 2024 · 2 repositories · arXiv:2408.14812
-
MROVSeg: Breaking the Resolution Curse of Vision-Language Models in Open-Vocabulary Image Segmentation 27 Aug 2024 · 0 repositories · arXiv:2408.14776
-
The Benefits of Balance: From Information Projections to Variance Reduction 27 Aug 2024 · 0 repositories · arXiv:2408.15065
-
Explaining Vision-Language Similarities in Dual Encoders with Feature-Pair Attributions 26 Aug 2024 · 0 repositories · arXiv:2408.14153
-
Nemesis: Normalizing the Soft-prompt Vectors of Vision-Language Models 26 Aug 2024 · 1 repository · arXiv:2408.13979Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 3 unverified (of 6 harvested samples) · 3 pointer-only (licence)
-
Smart Multi-Modal Search: Contextual Sparse and Dense Embedding Integration in Adobe Express 26 Aug 2024 · 0 repositories · arXiv:2408.14698
-
Social perception of faces in a vision-language model 26 Aug 2024 · 1 repository · arXiv:2408.14435Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 9 harvested samples)
-
SwiftBrush v2: Make Your One-step Diffusion Model Better Than Its Teacher 26 Aug 2024 · 1 repository · arXiv:2408.14176Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
LowCLIP: Adapting the CLIP Model Architecture for Low-Resource Languages in Multimodal Image Retrieval Task 25 Aug 2024 · 0 repositories · arXiv:2408.13909
-
Towards Completeness: A Generalizable Action Proposal Generator for Zero-Shot Temporal Action Localization 25 Aug 2024 · 1 repository · arXiv:2408.13777
-
EAViT: External Attention Vision Transformer for Audio Classification 23 Aug 2024 · 0 repositories · arXiv:2408.13201