Methods › Computer Vision › Vision and Language Pre-Trained Models › CLIP › Papers, page 18
Contrastive Language-Image Pre-training
CLIP
Papers archive 2025-07-28
archive papers tagged: 3,094 · with a code link: 1,617 · where Syntology ran a sample: 649 (554 with a run with no instrument failure, 95 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (649 of 3,094 tagged: 554 with a run with no instrument failure, 95 where every run was a failure of Syntology's instrument)
Page 18 of 31: papers 1,701 to 1,800 of 3,094, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Do Vision and Language Encoders Represent the World Similarly? 10 Jan 2024 · 1 repository · arXiv:2401.05224Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
SnapCap: Efficient Snapshot Compressive Video Captioning 10 Jan 2024 · 0 repositories · arXiv:2401.04903
-
Towards Online Continuous Sign Language Recognition and Translation 10 Jan 2024 · 1 repository · arXiv:2401.05336Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Pre-trained Model Guided Fine-Tuning for Zero-Shot Adversarial Robustness 9 Jan 2024 · 1 repository · arXiv:2401.04350Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 10 harvested samples) · 10 pointer-only (licence)
-
Benchmarking PathCLIP for Pathology Image Analysis 5 Jan 2024 · 0 repositories · arXiv:2401.02651
-
Denoising Vision Transformers 5 Jan 2024 · 1 repository · arXiv:2401.02957Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 1 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 7 harvested samples) · 2 pointer-only (licence)
-
Latte: Latent Diffusion Transformer for Video Generation 5 Jan 2024 · 4 repositories · arXiv:2401.03048Syntology official (archive's flag): 12 ran · 12 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 1 honoured, 0 violated, 9 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 13 harvested samples)
-
Open-Vocabulary SAM: Segment and Recognize Twenty-thousand Classes Interactively 5 Jan 2024 · 1 repository · arXiv:2401.02955Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 1 violated, 4 with no contract checked; 2 where Syntology's instrument failed) · 4 unverified (of 11 harvested samples) · 11 pointer-only (licence)
-
3D Open-Vocabulary Panoptic Segmentation with 2D-3D Vision-Language Distillation 4 Jan 2024 · 0 repositories · arXiv:2401.02402
-
A Dataset and Benchmark for Copyright Infringement Unlearning from Text-to-Image Diffusion Models 4 Jan 2024 · 1 repository · arXiv:2403.12052
-
ChangeCLIP: Remote sensing change detection with multimodal vision-language representation learning 4 Jan 2024 · 1 repository
-
Improved Zero-Shot Classification by Adapting VLMs with Text Descriptions 4 Jan 2024 · 1 repository · arXiv:2401.02460Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
Learning to Prompt with Text Only Supervision for Vision-Language Models 4 Jan 2024 · 1 repository · arXiv:2401.02418Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 3 where Syntology's instrument failed) · 2 unverified (of 6 harvested samples) · 3 pointer-only (licence)
-
Mining Fine-Grained Image-Text Alignment for Zero-Shot Captioning via Text-Only Training 4 Jan 2024 · 1 repository · arXiv:2401.02347Syntology official (archive's flag): 2 ran · 3 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 2 where Syntology's instrument failed) · 5 unverified (of 8 harvested samples) · 8 pointer-only (licence)
-
Prompt Decoupling for Text-to-Image Person Re-identification 4 Jan 2024 · 0 repositories · arXiv:2401.02173
-
SyCoCa: Symmetrizing Contrastive Captioners with Attentive Masking for Multimodal Alignment 4 Jan 2024 · 0 repositories · arXiv:2401.02137
-
Few-shot Adaptation of Multi-modal Foundation Models: A Survey 3 Jan 2024 · 0 repositories · arXiv:2401.01736
-
Incorporating Geo-Diverse Knowledge into Prompting for Increased Geographical Robustness in Object Recognition 3 Jan 2024 · 0 repositories · arXiv:2401.01482
-
Learning Prompt with Distribution-Based Feature Replay for Few-Shot Class-Incremental Learning 3 Jan 2024 · 1 repository · arXiv:2401.01598
-
ColorizeDiffusion: Adjustable Sketch Colorization with Reference Image and Text 2 Jan 2024 · 2 repositories · arXiv:2401.01456
-
DialCLIP: Empowering CLIP as Multi-Modal Dialog Retriever 2 Jan 2024 · 0 repositories · arXiv:2401.01076
-
A Pedestrian is Worth One Prompt: Towards Language Guidance Person Re-Identification 1 Jan 2024 · 0 repositories
-
AM-RADIO: Agglomerative Vision Foundation Model Reduce All Domains Into One 1 Jan 2024 · 1 repository
-
Bayesian Exploration of Pre-trained Models for Low-shot Image Classification 1 Jan 2024 · 0 repositories
-
Bilateral Adaptation for Human-Object Interaction Detection with Occlusion-Robustness 1 Jan 2024 · 0 repositories
-
Brush2Prompt: Contextual Prompt Generator for Object Inpainting 1 Jan 2024 · 0 repositories
-
Building Vision-Language Models on Solid Foundations with Masked Distillation 1 Jan 2024 · 0 repositories
-
CLIP-Driven Open-Vocabulary 3D Scene Graph Generation via Cross-Modality Contrastive Learning 1 Jan 2024 · 0 repositories
-
DeIL: Direct-and-Inverse CLIP for Open-World Few-Shot Learning 1 Jan 2024 · 1 repository
-
DiG-IN: Diffusion Guidance for Investigating Networks - Uncovering Classifier Differences Neuron Visualisations and Visual Counterfactual Explanations 1 Jan 2024 · 1 repository
-
Disentangled Prompt Representation for Domain Generalization 1 Jan 2024 · 0 repositories
-
Distilling CLIP with Dual Guidance for Learning Discriminative Human Body Shape Representation 1 Jan 2024 · 0 repositories
-
Enhanced Motion-Text Alignment for Image-to-Video Transfer Learning 1 Jan 2024 · 0 repositories
-
Exploring Regional Clues in CLIP for Zero-Shot Semantic Segmentation 1 Jan 2024 · 1 repository
-
Improved Self-Training for Test-Time Adaptation 1 Jan 2024 · 1 repository
-
Language-only Training of Zero-shot Composed Image Retrieval 1 Jan 2024 · 1 repository
-
Learn to Rectify the Bias of CLIP for Unsupervised Semantic Segmentation 1 Jan 2024 · 1 repository
-
Learning to Segment Referred Objects from Narrated Egocentric Videos 1 Jan 2024 · 0 repositories
-
MAPLM: A Real-World Large-Scale Vision-Language Benchmark for Map and Traffic Scene Understanding 1 Jan 2024 · 1 repository
-
Multimodal Prompt Perceiver: Empower Adaptiveness Generalizability and Fidelity for All-in-One Image Restoration 1 Jan 2024 · 0 repositories
-
Point Segment and Count: A Generalized Framework for Object Counting 1 Jan 2024 · 1 repository
-
Prompt-Driven Referring Image Segmentation with Instance Contrasting 1 Jan 2024 · 0 repositories
-
Towards Efficient and Effective Text-to-Video Retrieval with Coarse-to-Fine Visual Representation Learning 1 Jan 2024 · 1 repository · arXiv:2401.00701
-
Transductive Zero-Shot and Few-Shot CLIP 1 Jan 2024 · 1 repository
-
Tune-An-Ellipse: CLIP Has Potential to Find What You Want 1 Jan 2024 · 1 repository
-
Unknown Prompt the only Lacuna: Unveiling CLIP's Potential for Open Domain Generalization 1 Jan 2024 · 1 repository
-
Unlocking the Potential of Pre-trained Vision Transformers for Few-Shot Semantic Segmentation through Relationship Descriptors 1 Jan 2024 · 1 repository
-
COMMA: Co-Articulated Multi-Modal Learning 30 Dec 2023 · 1 repository · arXiv:2401.00268
-
GazeCLIP: Towards Enhancing Gaze Estimation via Text Guidance 30 Dec 2023 · 0 repositories · arXiv:2401.00260
-
Leveraging Open-Vocabulary Diffusion to Camouflaged Instance Segmentation 29 Dec 2023 · 0 repositories · arXiv:2312.17505
-
Learning Vision from Models Rivals Learning Vision from Data 28 Dec 2023 · 2 repositories · arXiv:2312.17742Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 2 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 1 pointer-only (licence)
-
MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices 28 Dec 2023 · 1 repository · arXiv:2312.16886
-
TinyGPT-V: Efficient Multimodal Large Language Model via Small Backbones 28 Dec 2023 · 2 repositories · arXiv:2312.16862Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 6 where Syntology's instrument failed) · 0 unverified (of 7 harvested samples)
-
Visual Explanations of Image-Text Representations via Multi-Modal Information Bottleneck Attribution 28 Dec 2023 · 1 repository · arXiv:2312.17174
-
Forgery-aware Adaptive Transformer for Generalizable Synthetic Image Detection 27 Dec 2023 · 2 repositories · arXiv:2312.16649Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 5 where Syntology's instrument failed) · 1 unverified (of 7 harvested samples) · 4 pointer-only (licence)
-
VLCounter: Text-aware Visual Representation for Zero-Shot Object Counting 27 Dec 2023 · 1 repository · arXiv:2312.16580
-
HarmonyView: Harmonizing Consistency and Diversity in One-Image-to-3D 26 Dec 2023 · 1 repository · arXiv:2312.15980Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
LangSplat: 3D Language Gaussian Splatting 26 Dec 2023 · 1 repository · arXiv:2312.16084Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
APTv2: Benchmarking Animal Pose Estimation and Tracking with a Large-scale Dataset and Beyond 25 Dec 2023 · 1 repository · arXiv:2312.15612
-
BrainVis: Exploring the Bridge between Brain and Visual Signals via Image Reconstruction 22 Dec 2023 · 1 repository · arXiv:2312.14871
-
FM-OV3D: Foundation Model-based Cross-modal Knowledge Blending for Open-Vocabulary 3D Detection 22 Dec 2023 · 0 repositories · arXiv:2312.14465
-
Leveraging Habitat Information for Fine-grained Bird Identification 22 Dec 2023 · 0 repositories · arXiv:2312.14999
-
Plan, Posture and Go: Towards Open-World Text-to-Motion Generation 22 Dec 2023 · 0 repositories · arXiv:2312.14828
-
Unveiling Backbone Effects in CLIP: Exploring Representational Synergies and Variances 22 Dec 2023 · 0 repositories · arXiv:2312.14400
-
Multi-Sentence Grounding for Long-term Instructional Video 21 Dec 2023 · 0 repositories · arXiv:2312.14055
-
Diff-Oracle: Deciphering Oracle Bone Scripts with Controllable Diffusion Model 21 Dec 2023 · 0 repositories · arXiv:2312.13631
-
Parrot Captions Teach CLIP to Spot Text 21 Dec 2023 · 1 repository · arXiv:2312.14232
-
Weakly Supervised Semantic Segmentation for Driving Scenes 21 Dec 2023 · 1 repository · arXiv:2312.13646
-
DVIS++: Improved Decoupled Framework for Universal Video Segmentation 20 Dec 2023 · 1 repository · arXiv:2312.13305
-
Mutual-modality Adversarial Attack with Semantic Perturbation 20 Dec 2023 · 0 repositories · arXiv:2312.12768
-
Spectral Prompt Tuning:Unveiling Unseen Classes for Zero-Shot Semantic Segmentation 20 Dec 2023 · 1 repository · arXiv:2312.12754Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 1 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
TagCLIP: A Local-to-Global Framework to Enhance Open-Vocabulary Multi-Label Classification of CLIP Without Training 20 Dec 2023 · 1 repository · arXiv:2312.12828
-
CLIP-DINOiser: Teaching CLIP a few DINO tricks for open-vocabulary semantic segmentation 19 Dec 2023 · 1 repository · arXiv:2312.12359Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Open Vocabulary Semantic Scene Sketch Understanding 18 Dec 2023 · 0 repositories · arXiv:2312.12463
-
Understanding the Multi-modal Prompts of the Pre-trained Vision-Language Model 18 Dec 2023 · 0 repositories · arXiv:2312.11570
-
CEIR: Concept-based Explainable Image Representation Learning 17 Dec 2023 · 0 repositories · arXiv:2312.10747
-
Pedestrian Attribute Recognition via CLIP based Prompt Vision-Language Fusion 17 Dec 2023 · 2 repositories · arXiv:2312.10692
-
SAI3D: Segment Any Instance in 3D Scenes 17 Dec 2023 · 0 repositories · arXiv:2312.11557
-
StarVector: Generating Scalable Vector Graphics Code from Images and Text 17 Dec 2023 · 1 repository · arXiv:2312.11556
-
CLIPSyntel: CLIP and LLM Synergy for Multimodal Question Summarization in Healthcare 16 Dec 2023 · 1 repository · arXiv:2312.11541
-
Learning Interpretable Queries for Explainable Image Classification with Information Pursuit 16 Dec 2023 · 0 repositories · arXiv:2312.11548
-
RetailKLIP : Finetuning OpenCLIP backbone using metric learning on a single GPU for Zero-shot retail product image classification 16 Dec 2023 · 0 repositories · arXiv:2312.10282
-
Shot2Story20K: A New Benchmark for Comprehensive Understanding of Multi-shot Videos 16 Dec 2023 · 1 repository · arXiv:2312.10300Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 7 harvested samples) · 7 pointer-only (licence)
-
Simple Image-level Classification Improves Open-vocabulary Object Detection 16 Dec 2023 · 1 repository · arXiv:2312.10439
-
Collaborating Foundation Models for Domain Generalized Semantic Segmentation 15 Dec 2023 · 1 repository · arXiv:2312.09788Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 7 harvested samples)
-
Data-Efficient Multimodal Fusion on a Single GPU 15 Dec 2023 · 2 repositories · arXiv:2312.10144
-
Osprey: Pixel Understanding with Visual Instruction Tuning 15 Dec 2023 · 2 repositories · arXiv:2312.10032Syntology official (archive's flag): 8 ran · 10 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 1 violated, 7 with no contract checked; 2 where Syntology's instrument failed) · 4 unverified (of 14 harvested samples) · 6 pointer-only (licence)
-
Structural Information Guided Multimodal Pre-training for Vehicle-centric Perception 15 Dec 2023 · 1 repository · arXiv:2312.09812
-
TAB: Text-Align Anomaly Backbone Model for Industrial Inspection Tasks 15 Dec 2023 · 0 repositories · arXiv:2312.09480
-
Toward Deep Drum Source Separation 15 Dec 2023 · 1 repository · arXiv:2312.09663
-
A Picture is Worth More Than 77 Text Tokens: Evaluating CLIP-Style Models on Dense Captions 14 Dec 2023 · 1 repository · arXiv:2312.08578Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 12 harvested samples) · 12 pointer-only (licence)
-
CLIP-guided Federated Learning on Heterogeneous and Long-Tailed Data 14 Dec 2023 · 1 repository · arXiv:2312.08648
-
Improving Cross-modal Alignment with Synthetic Pairs for Text-only Image Captioning 14 Dec 2023 · 0 repositories · arXiv:2312.08865
-
MmAP : Multi-modal Alignment Prompt for Cross-domain Multi-task Learning 14 Dec 2023 · 0 repositories · arXiv:2312.08636
-
OMG: Towards Open-vocabulary Motion Generation via Mixture of Controllers 14 Dec 2023 · 0 repositories · arXiv:2312.08985
-
Tokenize Anything via Prompting 14 Dec 2023 · 1 repository · arXiv:2312.09128Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 6 harvested samples)
-
Vision-Language Models as a Source of Rewards 14 Dec 2023 · 0 repositories · arXiv:2312.09187
-
Clockwork Diffusion: Efficient Generation With Model-Step Distillation 13 Dec 2023 · 1 repository · arXiv:2312.08128
-
EZ-CLIP: Efficient Zeroshot Video Action Recognition 13 Dec 2023 · 1 repository · arXiv:2312.08010Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 5 where Syntology's instrument failed) · 4 unverified (of 13 harvested samples) · 4 pointer-only (licence)
-
LAMM: Label Alignment for Multi-Modal Prompt Learning 13 Dec 2023 · 1 repository · arXiv:2312.08212