Methods › Computer Vision › Vision and Language Pre-Trained Models › CLIP › Papers, page 8
Contrastive Language-Image Pre-training
CLIP
Papers archive 2025-07-28
archive papers tagged: 3,094 · with a code link: 1,617 · where Syntology ran a sample: 649 (554 with a run with no instrument failure, 95 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (649 of 3,094 tagged: 554 with a run with no instrument failure, 95 where every run was a failure of Syntology's instrument)
Page 8 of 31: papers 701 to 800 of 3,094, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
DeDe: Detecting Backdoor Samples for SSL Encoders via Decoders 25 Nov 2024 · 1 repository · arXiv:2411.16154
-
ENCLIP: Ensembling and Clustering-Based Contrastive Language-Image Pretraining for Fashion Multimodal Search with Limited Data and Low-Quality Images 25 Nov 2024 · 0 repositories · arXiv:2411.16096
-
Factorized Visual Tokenization and Generation 25 Nov 2024 · 0 repositories · arXiv:2411.16681
-
Seq2Time: Sequential Knowledge Transfer for Video LLM Temporal Grounding 25 Nov 2024 · 0 repositories · arXiv:2411.16932
-
Soft-TransFormers for Continual Learning 25 Nov 2024 · 1 repository · arXiv:2411.16073
-
Style-Pro: Style-Guided Prompt Learning for Generalizable Vision-Language Models 25 Nov 2024 · 0 repositories · arXiv:2411.16018
-
LibraGrad: Balancing Gradient Flow for Universally Better Vision Transformer Attributions 24 Nov 2024 · 1 repository · arXiv:2411.16760
-
Modality Alignment Meets Federated Broadcasting 24 Nov 2024 · 0 repositories · arXiv:2411.15837
-
ResCLIP: Residual Attention for Training-free Dense Vision-language Inference 24 Nov 2024 · 1 repository · arXiv:2411.15851
-
Self-Calibrated CLIP for Training-Free Open-Vocabulary Segmentation 24 Nov 2024 · 1 repository · arXiv:2411.15869Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 3 where Syntology's instrument failed) · 2 unverified (of 7 harvested samples) · 2 pointer-only (licence)
-
MUNBa: Machine Unlearning via Nash Bargaining 23 Nov 2024 · 1 repository · arXiv:2411.15537Syntology official (archive's flag): 2 ran · 5 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 2 honoured, 2 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
Adversarial Prompt Distillation for Vision-Language Models 22 Nov 2024 · 0 repositories · arXiv:2411.15244
-
Detecting Visual Triggers in Cannabis Imagery: A CLIP-Based Multi-Labeling Framework with Local-Global Aggregation 22 Nov 2024 · 0 repositories · arXiv:2412.08648
-
Effective SAM Combination for Open-Vocabulary Semantic Segmentation 22 Nov 2024 · 0 repositories · arXiv:2411.14723
-
FloAt: Flow Warping of Self-Attention for Clothing Animation Generation 22 Nov 2024 · 0 repositories · arXiv:2411.15028
-
Open-Vocabulary Online Semantic Mapping for SLAM 22 Nov 2024 · 1 repository · arXiv:2411.15043
-
Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers 22 Nov 2024 · 0 repositories · arXiv:2411.14789
-
WildLMa: Long Horizon Loco-Manipulation in the Wild 22 Nov 2024 · 0 repositories · arXiv:2411.15131
-
BiomedCoOp: Learning to Prompt for Biomedical Vision-Language Models 21 Nov 2024 · 1 repository · arXiv:2411.15232
-
CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation 21 Nov 2024 · 1 repository · arXiv:2411.13836
-
FoPru: Focal Pruning for Efficient Large Vision-Language Models 21 Nov 2024 · 0 repositories · arXiv:2411.14164
-
Multimodal Autoregressive Pre-training of Large Vision Encoders 21 Nov 2024 · 1 repository · arXiv:2411.14402Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Hints of Prompt: Enhancing Visual Representation for Multimodal LLMs in Autonomous Driving 20 Nov 2024 · 0 repositories · arXiv:2411.13076
-
REDUCIO! Generating 1024×1024 Video within 16 Seconds using Extremely Compressed Motion Latents 20 Nov 2024 · 1 repository · arXiv:2411.13552Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
TAPT: Test-Time Adversarial Prompt Tuning for Robust Inference in Vision-Language Models 20 Nov 2024 · 0 repositories · arXiv:2411.13136
-
ViSTa Dataset: Do vision-language models understand sequential tasks? 20 Nov 2024 · 1 repository · arXiv:2411.13211
-
HyperGAN-CLIP: A Unified Framework for Domain Adaptation, Image Synthesis and Manipulation 19 Nov 2024 · 1 repository · arXiv:2411.12832
-
Joint Vision-Language Social Bias Removal for CLIP 19 Nov 2024 · 1 repository · arXiv:2411.12785
-
FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training 18 Nov 2024 · 1 repository · arXiv:2411.11927
-
ITACLIP: Boosting Training-Free Semantic Segmentation with Image, Text, and Architectural Enhancements 18 Nov 2024 · 1 repository · arXiv:2411.12044
-
Teaching Video Diffusion Model with Latent Physical Phenomenon Knowledge 18 Nov 2024 · 0 repositories · arXiv:2411.11343
-
Text-guided Zero-Shot Object Localization 18 Nov 2024 · 0 repositories · arXiv:2411.11357
-
Unveiling the Hidden: Online Vectorized HD Map Construction with Clip-Level Token Interaction and Propagation 17 Nov 2024 · 0 repositories · arXiv:2411.11002
-
MpoxVLM: A Vision-Language Model for Diagnosing Skin Lesions from Mpox Virus Infection 16 Nov 2024 · 1 repository · arXiv:2411.10888
-
CorrCLIP: Reconstructing Correlations in CLIP with Off-the-Shelf Foundation Models for Open-Vocabulary Semantic Segmentation 15 Nov 2024 · 1 repository · arXiv:2411.10086Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 9 harvested samples) · 9 pointer-only (licence)
-
Harnessing Vision Foundation Models for High-Performance, Training-Free Open Vocabulary Segmentation 14 Nov 2024 · 1 repository · arXiv:2411.09219
-
SCAN: Bootstrapping Contrastive Pre-training for Data Efficiency 14 Nov 2024 · 1 repository · arXiv:2411.09126
-
AstroM³: A self-supervised multimodal model for astronomy 13 Nov 2024 · 0 repositories · arXiv:2411.08842
-
Measuring similarity between embedding spaces using induced neighborhood graphs 13 Nov 2024 · 0 repositories · arXiv:2411.08687
-
Aligning Visual Contrastive learning models via Preference Optimization 12 Nov 2024 · 1 repository · arXiv:2411.08923
-
Contrastive Language Prompting to Ease False Positives in Medical Anomaly Detection 12 Nov 2024 · 1 repository · arXiv:2411.07546
-
Robust Fine-tuning of Zero-shot Models via Variance Reduction 11 Nov 2024 · 1 repository · arXiv:2411.06966Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 6 where Syntology's instrument failed) · 4 unverified (of 11 harvested samples) · 11 pointer-only (licence)
-
UMFC: Unsupervised Multi-Domain Feature Calibration for Vision-Language Models 11 Nov 2024 · 1 repository · arXiv:2411.06921Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Aquila-plus: Prompt-Driven Visual-Language Models for Pixel-Level Remote Sensing Image Understanding 9 Nov 2024 · 0 repositories · arXiv:2411.06142
-
ViTOC: Vision Transformer and Object-aware Captioner 9 Nov 2024 · 0 repositories · arXiv:2411.07265
-
Enhancing Visual Classification using Comparative Descriptors 8 Nov 2024 · 1 repository · arXiv:2411.05357
-
Integrating Object Detection Modality into Visual Language Model for Enhanced Autonomous Driving Agent 8 Nov 2024 · 0 repositories · arXiv:2411.05898
-
Image Understanding Makes for A Good Tokenizer for Image Generation 7 Nov 2024 · 1 repository · arXiv:2411.04406
-
In the Era of Prompt Learning with Vision-Language Models 7 Nov 2024 · 0 repositories · arXiv:2411.04892
-
LLM2CLIP: Powerful Language Model Unlocks Richer Visual Representation 7 Nov 2024 · 1 repository · arXiv:2411.04997
-
On Erroneous Agreements of CLIP Image Embeddings 7 Nov 2024 · 1 repository · arXiv:2411.05195Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 1 violated, 4 with no contract checked; 2 where Syntology's instrument failed) · 4 unverified (of 11 harvested samples) · 11 pointer-only (licence)
-
Customized Multiple Clustering via Multi-Modal Subspace Proxy Learning 6 Nov 2024 · 1 repository · arXiv:2411.03978Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 2 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
Textual Decomposition Then Sub-motion-space Scattering for Open-Vocabulary Motion Generation 6 Nov 2024 · 0 repositories · arXiv:2411.04079
-
Classification Done Right for Vision-Language Pre-Training 5 Nov 2024 · 1 repository · arXiv:2411.03313Syntology official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 5 where Syntology's instrument failed) · 8 unverified (of 18 harvested samples)
-
GraphVL: Graph-Enhanced Semantic Modeling via Vision-Language Models for Generalized Class Discovery 4 Nov 2024 · 0 repositories · arXiv:2411.02074
-
MM-Embed: Universal Multimodal Retrieval with Multimodal LLMs 4 Nov 2024 · 0 repositories · arXiv:2411.02571
-
PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance 4 Nov 2024 · 1 repository · arXiv:2411.02327Syntology official (archive's flag): 4 ran · 5 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 1 violated, 3 with no contract checked; 1 where Syntology's instrument failed) · 4 unverified (of 9 harvested samples) · 3 pointer-only (licence)
-
TripletCLIP: Improving Compositional Reasoning of CLIP via Synthetic Vision-Language Negatives 4 Nov 2024 · 0 repositories · arXiv:2411.02545
-
Finding NeMo: Negative-mined Mosaic Augmentation for Referring Image Segmentation 3 Nov 2024 · 0 repositories · arXiv:2411.01494
-
B-cosification: Transforming Deep Neural Networks to be Inherently Interpretable 1 Nov 2024 · 1 repository · arXiv:2411.00715Syntology official (archive's flag): 8 ran · 12 ran (of which 2 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 3 where Syntology's instrument failed) · 1 unverified (of 13 harvested samples)
-
Contrasting with Symile: Simple Model-Agnostic Representation Learning for Unlimited Modalities 1 Nov 2024 · 1 repository · arXiv:2411.01053Syntology official (archive's flag): 2 ran · 2 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Identifying Implicit Social Biases in Vision-Language Models 1 Nov 2024 · 0 repositories · arXiv:2411.00997
-
StyleTex: Style Image-Guided Texture Generation for 3D Models 1 Nov 2024 · 0 repositories · arXiv:2411.00399
-
Unified Generative and Discriminative Training for Multi-modal Large Language Models 1 Nov 2024 · 0 repositories · arXiv:2411.00304
-
Aggregate-and-Adapt Natural Language Prompts for Downstream Generalization of CLIP 31 Oct 2024 · 0 repositories · arXiv:2410.23698
-
An Individual Identity-Driven Framework for Animal Re-Identification 30 Oct 2024 · 1 repository · arXiv:2410.22927
-
CLIPErase: Efficient Unlearning of Visual-Textual Associations in CLIP 30 Oct 2024 · 0 repositories · arXiv:2410.23330
-
Multilingual Vision-Language Pre-training for the Remote Sensing Domain 30 Oct 2024 · 1 repository · arXiv:2410.23370
-
Active Learning for Vision-Language Models 29 Oct 2024 · 0 repositories · arXiv:2410.22187
-
HairDiffusion: Vivid Multi-Colored Hair Editing via Latent Diffusion 29 Oct 2024 · 0 repositories · arXiv:2410.21789
-
Multi-Class Textual-Inversion Secretly Yields a Semantic-Agnostic Classifier 29 Oct 2024 · 1 repository · arXiv:2410.22317
-
Semi-Supervised Self-Learning Enhanced Music Emotion Recognition 29 Oct 2024 · 0 repositories · arXiv:2410.21897
-
Text-Guided Attention is All You Need for Zero-Shot Robustness in Vision-Language Models 29 Oct 2024 · 1 repository · arXiv:2410.21802Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 6 unverified (of 12 harvested samples) · 12 pointer-only (licence)
-
Attention Overlap Is Responsible for The Entity Missing Problem in Text-to-image Diffusion Models! 28 Oct 2024 · 0 repositories · arXiv:2410.20972
-
David and Goliath: Small One-step Model Beats Large Diffusion with Score Post-training 28 Oct 2024 · 1 repository · arXiv:2410.20898
-
FewVS: A Vision-Semantics Integration Framework for Few-Shot Image Classification 28 Oct 2024 · 1 repository
-
Flexible Natural Language-Based Image Data Downlink Prioritization for Nanosatellites 28 Oct 2024 · 1 repository
-
Semantic Editing Increment Benefits Zero-Shot Composed Image Retrieval 28 Oct 2024 · 2 repositories
-
R-LLaVA: Improving Med-VQA Understanding through Visual Region of Interest 27 Oct 2024 · 0 repositories · arXiv:2410.20327
-
You Never Know: Quantization Induces Inconsistent Biases in Vision-Language Foundation Models 26 Oct 2024 · 0 repositories · arXiv:2410.20265
-
Enhancing Zero-Shot Vision Models by Label-Free Prompt Distribution Learning and Bias Correcting 25 Oct 2024 · 0 repositories · arXiv:2410.19294Syntology 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
NeuroClips: Towards High-fidelity and Smooth fMRI-to-Video Reconstruction 25 Oct 2024 · 1 repository · arXiv:2410.19452Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
ConceptDrift: Uncovering Biases through the Lens of Foundation Models 24 Oct 2024 · 0 repositories · arXiv:2410.18970
-
WAFFLE: Finetuning Multi-Modal Model for Automated Front-End Development 24 Oct 2024 · 1 repository · arXiv:2410.18362
-
Backdoor in Seconds: Unlocking Vulnerabilities in Large Pre-trained Models via Model Editing 23 Oct 2024 · 0 repositories · arXiv:2410.18267
-
EntityCLIP: Entity-Centric Image-Text Matching via Multimodal Attentive Contrastive Learning 23 Oct 2024 · 0 repositories · arXiv:2410.17810
-
Are Visual-Language Models Effective in Action Recognition? A Comparative Study 22 Oct 2024 · 0 repositories · arXiv:2410.17149
-
Benchmarking Large Language Models for Image Classification of Marine Mammals 22 Oct 2024 · 1 repository · arXiv:2410.19848
-
FairLoRA: Unpacking Bias Mitigation in Vision Models with Fairness-Driven Low-Rank Adaptation 22 Oct 2024 · 0 repositories · arXiv:2410.17358
-
An Efficient System for Automatic Map Storytelling -- A Case Study on Historical Maps 21 Oct 2024 · 1 repository · arXiv:2410.15780
-
In Search of the Successful Interpolation: On the Role of Sharpness in CLIP Generalization 21 Oct 2024 · 1 repository · arXiv:2410.16476
-
Visual Motif Identification: Elaboration of a Curated Comparative Dataset and Classification Methods 21 Oct 2024 · 0 repositories · arXiv:2410.15866
-
BoostAdapter: Improving Vision-Language Test-Time Adaptation via Regional Bootstrapping 20 Oct 2024 · 1 repository · arXiv:2410.15430Syntology official (archive's flag): 4 ran · 5 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 5 where Syntology's instrument failed) · 5 unverified (of 10 harvested samples) · 1 pointer-only (licence)
-
IPO: Interpretable Prompt Optimization for Vision-Language Models 20 Oct 2024 · 1 repository · arXiv:2410.15397Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 6 where Syntology's instrument failed) · 2 unverified (of 11 harvested samples) · 11 pointer-only (licence)
-
LoRA-IR: Taming Low-Rank Experts for Efficient All-in-One Image Restoration 20 Oct 2024 · 1 repository · arXiv:2410.15385Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 8 harvested samples) · 2 pointer-only (licence)
-
Open-vocabulary vs. Closed-set: Best Practice for Few-shot Object Detection Considering Text Describability 20 Oct 2024 · 1 repository · arXiv:2410.15315
-
Scene Graph Generation with Role-Playing Large Language Models 20 Oct 2024 · 1 repository · arXiv:2410.15364Syntology official (archive's flag): 2 ran · 2 ran (of which 2 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified; every one of the 2 samples that ran constructed an object rather than computing a result (of 3 harvested samples)
-
BYOCL: Build Your Own Consistent Latent with Hierarchical Representative Latent Clustering 19 Oct 2024 · 1 repository · arXiv:2410.15060
-
CLIPtortionist: Zero-shot Text-driven Deformation for Manufactured 3D Shapes 19 Oct 2024 · 0 repositories · arXiv:2410.15199
-
Visual Navigation of Digital Libraries: Retrieval and Classification of Images in the National Library of Norway's Digitised Book Collection 19 Oct 2024 · 1 repository · arXiv:2410.14969