Methods › Computer Vision › Vision and Language Pre-Trained Models › CLIP › Papers, page 11
Contrastive Language-Image Pre-training
CLIP
Papers archive 2025-07-28
archive papers tagged: 3,094 · with a code link: 1,617 · where Syntology ran a sample: 649 (554 with a run with no instrument failure, 95 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (649 of 3,094 tagged: 554 with a run with no instrument failure, 95 where every run was a failure of Syntology's instrument)
Page 11 of 31: papers 1,001 to 1,100 of 3,094, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Image Segmentation in Foundation Model Era: A Survey 23 Aug 2024 · 1 repository · arXiv:2408.12957
-
La-SoftMoE CLIP for Unified Physical-Digital Face Attack Detection 23 Aug 2024 · 0 repositories · arXiv:2408.12793
-
Online Zero-Shot Classification with CLIP 23 Aug 2024 · 1 repository · arXiv:2408.13320Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample)
-
QD-VMR: Query Debiasing with Contextual Understanding Enhancement for Video Moment Retrieval 23 Aug 2024 · 0 repositories · arXiv:2408.12981
-
Adapt CLIP as Aggregation Instructor for Image Dehazing 22 Aug 2024 · 0 repositories · arXiv:2408.12317
-
Visual Verity in AI-Generated Imagery: Computational Metrics and Human-Centric Analysis 22 Aug 2024 · 0 repositories · arXiv:2408.12762
-
EAGLE: Elevating Geometric Reasoning through LLM-empowered Visual Instruction Tuning 21 Aug 2024 · 0 repositories · arXiv:2408.11397
-
Enabling Small Models for Zero-Shot Selection and Reuse through Model Label Learning 21 Aug 2024 · 0 repositories · arXiv:2408.11449
-
Interpretable Long-term Action Quality Assessment 21 Aug 2024 · 1 repository · arXiv:2408.11687
-
SEA: Supervised Embedding Alignment for Token-Level Visual-Textual Integration in MLLMs 21 Aug 2024 · 0 repositories · arXiv:2408.11813
-
Generalizable Facial Expression Recognition 20 Aug 2024 · 1 repository · arXiv:2408.10614Syntology official (archive's flag): 5 ran · 5 ran (of which 2 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
Is the Lecture Engaging for Learning? Lecture Voice Sentiment Analysis for Knowledge Graph-Supported Intelligent Lecturing Assistant (ILA) System 20 Aug 2024 · 1 repository · arXiv:2408.10492
-
MUSE: Mamba is Efficient Multi-scale Learner for Text-video Retrieval 20 Aug 2024 · 1 repository · arXiv:2408.10575
-
A Unified Framework for Iris Anti-Spoofing: Introducing IrisGeneral Dataset and Masked-MoE Method 19 Aug 2024 · 0 repositories · arXiv:2408.09752
-
Boosting Open-Domain Continual Learning via Leveraging Intra-domain Category-aware Prototype 19 Aug 2024 · 0 repositories · arXiv:2408.09984
-
C2P-CLIP: Injecting Category Common Prompt in CLIP to Enhance Generalization in Deepfake Detection 19 Aug 2024 · 1 repository · arXiv:2408.09647Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
Caption-Driven Explorations: Aligning Image and Text Embeddings through Human-Inspired Foveated Vision 19 Aug 2024 · 0 repositories · arXiv:2408.09948
-
CLIP-DPO: Vision-Language Models as a Source of Preference for Fixing Hallucinations in LVLMs 19 Aug 2024 · 0 repositories · arXiv:2408.10433
-
CLIPCleaner: Cleaning Noisy Labels with CLIP 19 Aug 2024 · 1 repository · arXiv:2408.10012
-
Cross-composition Feature Disentanglement for Compositional Zero-shot Learning 19 Aug 2024 · 0 repositories · arXiv:2408.09786
-
SANER: Annotation-free Societal Attribute Neutralizer for Debiasing CLIP 19 Aug 2024 · 0 repositories · arXiv:2408.10202
-
CLIP-CID: Efficient CLIP Distillation via Cluster-Instance Discrimination 18 Aug 2024 · 0 repositories · arXiv:2408.09441
-
DPA: Dual Prototypes Alignment for Unsupervised Adaptation of Vision-Language Models 16 Aug 2024 · 1 repository · arXiv:2408.08855
-
TextCAVs: Debugging vision models using text 16 Aug 2024 · 1 repository · arXiv:2408.08652
-
TEXTOC: Text-driven Object-Centric Style Transfer 16 Aug 2024 · 0 repositories · arXiv:2408.08461
-
Heavy Labels Out! Dataset Distillation with Label Space Lightening 15 Aug 2024 · 0 repositories · arXiv:2408.08201
-
Navigating Data Scarcity using Foundation Models: A Benchmark of Few-Shot and Zero-Shot Learning Approaches in Medical Imaging 15 Aug 2024 · 1 repository · arXiv:2408.08058
-
Dual-Domain CLIP-Assisted Residual Optimization Perception Model for Metal Artifact Reduction 14 Aug 2024 · 0 repositories · arXiv:2408.14342
-
ReCLIP++: Learn to Rectify the Bias of CLIP for Unsupervised Semantic Segmentation 13 Aug 2024 · 1 repository · arXiv:2408.06747Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
Visual Neural Decoding via Improved Visual-EEG Semantic Consistency 13 Aug 2024 · 0 repositories · arXiv:2408.06788
-
Freehand Sketch Generation from Mechanical Components 12 Aug 2024 · 1 repository · arXiv:2408.05966
-
3D-free meets 3D priors: Novel View Synthesis from a Single Image with Pretrained Diffusion Guidance 12 Aug 2024 · 0 repositories · arXiv:2408.06157
-
OmniCLIP: Adapting CLIP for Video Recognition with Spatial-Temporal Omni-Scale Feature Learning 12 Aug 2024 · 1 repository · arXiv:2408.06158
-
Prompt Recovery for Image Generation Models: A Comparative Study of Discrete Optimizers 12 Aug 2024 · 0 repositories · arXiv:2408.06502
-
Unseen No More: Unlocking the Potential of CLIP for Generative Zero-shot HOI Detection 12 Aug 2024 · 1 repository · arXiv:2408.05974
-
Decoder Pre-Training with only Text for Scene Text Recognition 11 Aug 2024 · 1 repository · arXiv:2408.05706
-
Efficient and Versatile Robust Fine-Tuning of Zero-shot Models 11 Aug 2024 · 0 repositories · arXiv:2408.05749
-
Hyperbolic Learning with Multimodal Large Language Models 9 Aug 2024 · 0 repositories · arXiv:2408.05097
-
ProxyCLIP: Proxy Attention Improves CLIP for Open-Vocabulary Segmentation 9 Aug 2024 · 1 repository · arXiv:2408.04883Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 1 violated, 3 with no contract checked; 2 where Syntology's instrument failed) · 3 unverified (of 9 harvested samples) · 9 pointer-only (licence)
-
Weak-Annotation of HAR Datasets using Vision Foundation Models 9 Aug 2024 · 1 repository · arXiv:2408.05169
-
ComKD-CLIP: Comprehensive Knowledge Distillation for Contrastive Language-Image Pre-traning Model 8 Aug 2024 · 0 repositories · arXiv:2408.04145
-
Ensemble everything everywhere: Multi-scale aggregation for adversarial robustness 8 Aug 2024 · 2 repositories · arXiv:2408.05446Syntology official: no sample here; runs from other or unrecorded repositories · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
LLDif: Diffusion Models for Low-light Emotion Recognition 8 Aug 2024 · 0 repositories · arXiv:2408.04235
-
ArtVLM: Attribute Recognition Through Vision-Based Prefix Language Modeling 7 Aug 2024 · 1 repository · arXiv:2408.04102
-
CLIP-based Point Cloud Classification via Point Cloud to Image Translation 7 Aug 2024 · 0 repositories · arXiv:2408.03545
-
How Well Can Vision Language Models See Image Details? 7 Aug 2024 · 0 repositories · arXiv:2408.03940
-
MoExtend: Tuning New Experts for Modality and Task Extension 7 Aug 2024 · 1 repository · arXiv:2408.03511Syntology official (archive's flag): 3 ran · 3 ran (of which 2 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples)
-
Teach CLIP to Develop a Number Sense for Ordinal Regression 7 Aug 2024 · 1 repository · arXiv:2408.03574
-
Vision-Language Guidance for LiDAR-based Unsupervised 3D Object Detection 7 Aug 2024 · 1 repository · arXiv:2408.03790
-
Explain via Any Concept: Concept Bottleneck Model with Open Vocabulary Concepts 5 Aug 2024 · 0 repositories · arXiv:2408.02265
-
Exploring Conditional Multi-Modal Prompts for Zero-shot HOI Detection 5 Aug 2024 · 1 repository · arXiv:2408.02484Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 2 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
Latent-INR: A Flexible Framework for Implicit Representations of Videos with Discriminative Semantics 5 Aug 2024 · 0 repositories · arXiv:2408.02672
-
Text Conditioned Symbolic Drumbeat Generation using Latent Diffusion Models 5 Aug 2024 · 1 repository · arXiv:2408.02711
-
Dataset Scale and Societal Consistency Mediate Facial Impression Bias in Vision-Language AI 4 Aug 2024 · 0 repositories · arXiv:2408.01959
-
ML-EAT: A Multilevel Embedding Association Test for Interpretable and Transparent Social Science 4 Aug 2024 · 1 repository · arXiv:2408.01966
-
AdvQDet: Detecting Query-Based Adversarial Attacks with Adversarial Contrastive Prompt Tuning 4 Aug 2024 · 1 repository · arXiv:2408.01978Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 1 violated, 3 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 7 harvested samples) · 7 pointer-only (licence)
-
AdaCBM: An Adaptive Concept Bottleneck Model for Explainable and Accurate Diagnosis 4 Aug 2024 · 1 repository · arXiv:2408.02001
-
SAT3D: Image-driven Semantic Attribute Transfer in 3D 3 Aug 2024 · 0 repositories · arXiv:2408.01664
-
Exploiting the Semantic Knowledge of Pre-trained Text-Encoders for Continual Learning 2 Aug 2024 · 1 repository · arXiv:2408.01076Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
VAR-CLIP: Text-to-Image Generator with Visual Auto-Regressive Modeling 2 Aug 2024 · 1 repository · arXiv:2408.01181Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 5 where Syntology's instrument failed) · 4 unverified (of 13 harvested samples) · 13 pointer-only (licence)
-
CLIP4Sketch: Enhancing Sketch to Mugshot Matching through Dataset Augmentation using Diffusion Models 2 Aug 2024 · 0 repositories · arXiv:2408.01233
-
Toward Automatic Relevance Judgment using Vision--Language Models for Image--Text Retrieval Evaluation 2 Aug 2024 · 0 repositories · arXiv:2408.01363
-
CIResDiff: A Clinically-Informed Residual Diffusion Model for Predicting Idiopathic Pulmonary Fibrosis Progression 1 Aug 2024 · 0 repositories · arXiv:2408.00938
-
A new approach for encoding code and assisting code understanding 1 Aug 2024 · 0 repositories · arXiv:2408.00521
-
Collaborative Vision-Text Representation Optimizing for Open-Vocabulary Segmentation 1 Aug 2024 · 1 repository · arXiv:2408.00744
-
Focus, Distinguish, and Prompt: Unleashing CLIP for Efficient and Flexible Scene Text Retrieval 1 Aug 2024 · 1 repository · arXiv:2408.00441Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
MOSAIC: Multimodal Multistakeholder-aware Visual Art Recommendation 31 Jul 2024 · 0 repositories · arXiv:2407.21758
-
EZSR: Event-based Zero-Shot Recognition 31 Jul 2024 · 0 repositories · arXiv:2407.21616
-
Generalized Out-of-Distribution Detection and Beyond in Vision Language Model Era: A Survey 31 Jul 2024 · 0 repositories · arXiv:2407.21794
-
MTA-CLIP: Language-Guided Semantic Segmentation with Mask-Text Alignment 31 Jul 2024 · 0 repositories · arXiv:2407.21654
-
Assessing Graphical Perception of Image Embedding Models using Channel Effectiveness 30 Jul 2024 · 0 repositories · arXiv:2407.20845
-
Bayesian Low-Rank LeArning (Bella): A Practical Approach to Bayesian Neural Networks 30 Jul 2024 · 1 repository · arXiv:2407.20891
-
Effectively Leveraging CLIP for Generating Situational Summaries of Images and Videos 30 Jul 2024 · 1 repository · arXiv:2407.20642
-
Image Re-Identification: Where Self-supervision Meets Vision-Language Learning 30 Jul 2024 · 1 repository · arXiv:2407.20647
-
Prompt-Driven Contrastive Learning for Transferable Adversarial Attacks 30 Jul 2024 · 0 repositories · arXiv:2407.20657
-
ActivityCLIP: Enhancing Group Activity Recognition by Mining Complementary Information from Text to Supplement Image Modality 29 Jul 2024 · 0 repositories · arXiv:2407.19820
-
Advancing Prompt Learning through an External Layer 29 Jul 2024 · 0 repositories · arXiv:2407.19674
-
Contrasting Deepfakes Diffusion via Contrastive Learning and Global-Local Similarities 29 Jul 2024 · 1 repository · arXiv:2407.20337Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 8 harvested samples)
-
Diffusion Feedback Helps CLIP See Better 29 Jul 2024 · 1 repository · arXiv:2407.20171Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 4 harvested samples)
-
Image-text matching for large-scale book collections 29 Jul 2024 · 1 repository · arXiv:2407.19812
-
MaskInversion: Localized Embeddings via Optimization of Explainability Maps 29 Jul 2024 · 0 repositories · arXiv:2407.20034
-
Faster Image2Video Generation: A Closer Look at CLIP Image Embedding's Impact on Spatio-Temporal Cross-Attentions 27 Jul 2024 · 0 repositories · arXiv:2407.19205
-
Adversarial Robustification via Text-to-Image Diffusion Models 26 Jul 2024 · 1 repository · arXiv:2407.18658Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 2 unverified (of 6 harvested samples) · 4 pointer-only (licence)
-
HICEScore: A Hierarchical Metric for Image Captioning Evaluation 26 Jul 2024 · 0 repositories · arXiv:2407.18589
-
𝕏-Sample Contrastive Loss: Improving Contrastive Learning with Sample Similarity Graphs 25 Jul 2024 · 0 repositories · arXiv:2407.18134
-
Unified Lexical Representation for Interpretable Visual-Language Alignment 25 Jul 2024 · 1 repository · arXiv:2407.17827
-
LangOcc: Self-Supervised Open Vocabulary Occupancy Estimation via Volume Rendering 24 Jul 2024 · 0 repositories · arXiv:2407.17310
-
Multi-label Cluster Discrimination for Visual Representation Learning 24 Jul 2024 · 1 repository · arXiv:2407.17331Syntology official (archive's flag): 7 ran · 7 ran (of which 7 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified; every one of the 7 samples that ran constructed an object rather than computing a result (of 11 harvested samples)
-
Selective Vision-Language Subspace Projection for Few-shot CLIP 24 Jul 2024 · 1 repository · arXiv:2407.16977
-
Unpaired Photo-realistic Image Deraining with Energy-informed Diffusion Model 24 Jul 2024 · 0 repositories · arXiv:2407.17193
-
Category-Extensible Out-of-Distribution Detection via Hierarchical Context Descriptions 23 Jul 2024 · 1 repository · arXiv:2407.16725Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 2 unverified (of 5 harvested samples) · 2 pointer-only (licence)
-
MLLM-CompBench: A Comparative Reasoning Benchmark for Multimodal LLMs 23 Jul 2024 · 1 repository · arXiv:2407.16837Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
SEDS: Semantically Enhanced Dual-Stream Encoder for Sign Language Retrieval 23 Jul 2024 · 1 repository · arXiv:2407.16394Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples)
-
VisMin: Visual Minimal-Change Understanding 23 Jul 2024 · 0 repositories · arXiv:2407.16772
-
AdaCLIP: Adapting CLIP with Hybrid Learnable Prompts for Zero-Shot Anomaly Detection 22 Jul 2024 · 1 repository · arXiv:2407.15795Syntology official (archive's flag): 20 ran · 20 ran (of which 10 constructed an object rather than computing a result; 16 with no instrument failure: 0 honoured, 1 violated, 15 with no contract checked; 4 where Syntology's instrument failed) · 8 unverified (of 28 harvested samples) · 4 pointer-only (licence)
-
CLIP with Generative Latent Replay: a Strong Baseline for Incremental Learning 22 Jul 2024 · 1 repository · arXiv:2407.15793Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 8 harvested samples) · 2 pointer-only (licence)
-
Reconstructing Training Data From Real World Models Trained with Transfer Learning 22 Jul 2024 · 0 repositories · arXiv:2407.15845
-
SAM2CLIP2SAM: Vision Language Model for Segmentation of 3D CT Scans for Covid-19 Detection 22 Jul 2024 · 0 repositories · arXiv:2407.15728
-
SLVideo: A Sign Language Video Moment Retrieval Framework 22 Jul 2024 · 0 repositories · arXiv:2407.15668
-
Assessing Brittleness of Image-Text Retrieval Benchmarks from Vision-Language Models Perspective 21 Jul 2024 · 0 repositories · arXiv:2407.15239