Methods › Computer Vision › Vision and Language Pre-Trained Models › CLIP › Papers, page 23
Contrastive Language-Image Pre-training
CLIP
Papers archive 2025-07-28
archive papers tagged: 3,094 · with a code link: 1,617 · where Syntology ran a sample: 649 (554 with a run with no instrument failure, 95 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (649 of 3,094 tagged: 554 with a run with no instrument failure, 95 where every run was a failure of Syntology's instrument)
Page 23 of 31: papers 2,201 to 2,300 of 3,094, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
In Defense of Clip-based Video Relation Detection 18 Jul 2023 · 0 repositories · arXiv:2307.08984
-
Multi-Stage Cable Routing through Hierarchical Imitation Learning 18 Jul 2023 · 0 repositories · arXiv:2307.08927
-
OnlineRefer: A Simple Online Baseline for Referring Video Object Segmentation 18 Jul 2023 · 1 repository · arXiv:2307.09356Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 1 honoured, 0 violated, 3 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 7 harvested samples) · 4 pointer-only (licence)
-
Text-guided Image Restoration and Semantic Enhancement for Text-to-Image Person Retrieval 18 Jul 2023 · 1 repository · arXiv:2307.09059
-
CLIP-Guided StyleGAN Inversion for Text-Driven Real Image Editing 17 Jul 2023 · 0 repositories · arXiv:2307.08397
-
Fast Adaptation with Bradley-Terry Preference Models in Text-To-Image Classification and Generation 15 Jul 2023 · 0 repositories · arXiv:2308.07929
-
Fine-grained Text-Video Retrieval with Frozen Image Encoders 14 Jul 2023 · 0 repositories · arXiv:2307.09972
-
Improving Zero-Shot Generalization for CLIP with Synthesized Prompts 14 Jul 2023 · 1 repository · arXiv:2307.07397Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
MMSD2.0: Towards a Reliable Multi-modal Sarcasm Detection System 14 Jul 2023 · 1 repository · arXiv:2307.07135
-
TALL: Thumbnail Layout for Deepfake Video Detection 14 Jul 2023 · 1 repository · arXiv:2307.07494Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 1 honoured, 0 violated, 5 with no contract checked; 2 where Syntology's instrument failed) · 2 unverified (of 10 harvested samples) · 1 pointer-only (licence)
-
AvatarFusion: Zero-shot Generation of Clothing-Decoupled 3D Avatars Using 2D Diffusion 13 Jul 2023 · 0 repositories · arXiv:2307.06526
-
Domain-Agnostic Tuning-Encoder for Fast Personalization of Text-To-Image Models 13 Jul 2023 · 0 repositories · arXiv:2307.06925
-
Leveraging Vision-Language Foundation Models for Fine-Grained Downstream Tasks 13 Jul 2023 · 1 repository · arXiv:2307.06795
-
Self-regulating Prompts: Foundational Model Adaptation without Forgetting 13 Jul 2023 · 2 repositories · arXiv:2307.06948Syntology official (archive's flag): 4 ran · 17 ran (of which 0 constructed an object rather than computing a result; 13 with no instrument failure: 0 honoured, 0 violated, 13 with no contract checked; 4 where Syntology's instrument failed) · 4 unverified (of 21 harvested samples) · 3 pointer-only (licence)
-
VELMA: Verbalization Embodiment of LLM Agents for Vision and Language Navigation in Street View 12 Jul 2023 · 1 repository · arXiv:2307.06082Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
MoP-CLIP: A Mixture of Prompt-Tuned CLIP Models for Domain Incremental Learning 11 Jul 2023 · 0 repositories · arXiv:2307.05707
-
PIGEON: Predicting Image Geolocations 11 Jul 2023 · 1 repository · arXiv:2307.05845Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Articulated 3D Head Avatar Generation using Text-to-Image Diffusion Models 10 Jul 2023 · 0 repositories · arXiv:2307.04859
-
CREPE: Learnable Prompting With CLIP Improves Visual Relationship Prediction 10 Jul 2023 · 1 repository · arXiv:2307.04838
-
Divide, Evaluate, and Refine: Evaluating and Improving Text-to-Image Alignment with Iterative VQA Feedback 10 Jul 2023 · 0 repositories · arXiv:2307.04749
-
Text Descriptions are Compressive and Invariant Representations for Visual Learning 10 Jul 2023 · 0 repositories · arXiv:2307.04317
-
mCLIP: Multilingual CLIP via Cross-lingual Transfer 10 Jul 2023 · 1 repository
-
Linear Alignment of Vision-language Models for Image Captioning 10 Jul 2023 · 1 repository · arXiv:2307.05591
-
Augmenters at SemEval-2023 Task 1: Enhancing CLIP in Handling Compositionality and Ambiguity for Zero-Shot Visual WSD through Prompt Augmentation and Text-To-Image Diffusion 9 Jul 2023 · 0 repositories · arXiv:2307.05564
-
CognitiveNet: Enriching Foundation Models with Emotions and Awareness 9 Jul 2023 · 0 repositories
-
Measuring the Success of Diffusion Models at Imitating Human Artists 8 Jul 2023 · 0 repositories · arXiv:2307.04028
-
Fooling Contrastive Language-Image Pre-trained Models with CLIPMasterPrints 7 Jul 2023 · 1 repository · arXiv:2307.03798
-
TBGC: Task-level Backbone-Oriented Gradient Clip for Multi-Task Foundation Model Learning 7 Jul 2023 · 0 repositories · arXiv:2307.03465
-
Deep Ensemble Learning with Frame Skipping for Face Anti-Spoofing 6 Jul 2023 · 3 repositories · arXiv:2307.02858
-
Proto-CLIP: Vision-Language Prototypical Network for Few-Shot Learning 6 Jul 2023 · 1 repository · arXiv:2307.03073
-
T-MARS: Improving Visual Representations by Circumventing Text Feature Learning 6 Jul 2023 · 1 repository · arXiv:2307.03132Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 6 harvested samples)
-
A ChatGPT Aided Explainable Framework for Zero-Shot Medical Image Diagnosis 5 Jul 2023 · 0 repositories · arXiv:2307.01981
-
Continual Learning in Open-vocabulary Classification with Complementary Memory Systems 4 Jul 2023 · 1 repository · arXiv:2307.01430
-
ClipSitu: Effectively Leveraging CLIP for Conditional Predictions in Situation Recognition 2 Jul 2023 · 1 repository · arXiv:2307.00586
-
ProbVLM: Probabilistic Adapter for Frozen Vision-Language Models 1 Jul 2023 · 1 repository · arXiv:2307.00398Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 3 where Syntology's instrument failed) · 7 unverified (of 15 harvested samples) · 2 pointer-only (licence)
-
CLIPAG: Towards Generator-Free Text-to-Image Generation 29 Jun 2023 · 0 repositories · arXiv:2306.16805
-
DreamDiffusion: Generating High-Quality Images from Brain EEG Signals 29 Jun 2023 · 1 repository · arXiv:2306.16934
-
Benchmarking Zero-Shot Recognition with Vision-Language Models: Challenges on Granularity and Specificity 28 Jun 2023 · 0 repositories · arXiv:2306.16048
-
ICSVR: Investigating Compositional and Syntactic Understanding in Video Retrieval Models 28 Jun 2023 · 2 repositories · arXiv:2306.16533
-
Pseudo-Labeling Enhanced by Privileged Information and Its Application to In Situ Sequencing Images 28 Jun 2023 · 0 repositories · arXiv:2306.15898
-
SpotEM: Efficient Video Search for Episodic Memory 28 Jun 2023 · 0 repositories · arXiv:2306.15850
-
Approximated Prompt Tuning for Vision-Language Pre-trained Models 27 Jun 2023 · 0 repositories · arXiv:2306.15706
-
CLIPA-v2: Scaling CLIP Training with 81.1% Zero-shot ImageNet Accuracy within a $10,000 Budget; An Extra $4,000 Unlocks 81.8% Accuracy 27 Jun 2023 · 2 repositories · arXiv:2306.15658
-
A Badminton Recognition and Tracking System Based on Context Multi-feature Fusion 26 Jun 2023 · 0 repositories · arXiv:2306.14492
-
Self-Supervised Image Captioning with CLIP 26 Jun 2023 · 0 repositories · arXiv:2306.15111
-
TCEIP: Text Condition Embedded Regression Network for Dental Implant Position Prediction 26 Jun 2023 · 0 repositories · arXiv:2306.14406
-
Addressing Cold Start Problem for End-to-end Automatic Speech Scoring 25 Jun 2023 · 0 repositories · arXiv:2306.14310
-
Learning-to-Rank Meets Language: Boosting Language-Driven Ordering Alignment for Ordinal Classification 24 Jun 2023 · 2 repositories · arXiv:2306.13856Syntology official (archive's flag): 11 ran · 18 ran (of which 0 constructed an object rather than computing a result; 11 with no instrument failure: 1 honoured, 0 violated, 10 with no contract checked; 7 where Syntology's instrument failed) · 5 unverified (of 23 harvested samples) · 5 pointer-only (licence)
-
Multimodal Search on Iconclass using Vision-Language Pre-Trained Models 23 Jun 2023 · 0 repositories · arXiv:2306.16529
-
TaCA: Upgrading Your Visual Foundation Model with Task-agnostic Compatible Adapter 22 Jun 2023 · 0 repositories · arXiv:2306.12642
-
Local 3D Editing via 3D Distillation of CLIP Knowledge 21 Jun 2023 · 0 repositories · arXiv:2306.12570
-
Mass-Producing Failures of Multimodal Systems with Language Models 21 Jun 2023 · 1 repository · arXiv:2306.12105Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
NeuroCLIP: Neuromorphic Data Understanding by CLIP and SNN 21 Jun 2023 · 1 repository · arXiv:2306.12073
-
Masking meets Supervision: A Strong Learning Alliance 20 Jun 2023 · 1 repository · arXiv:2306.11339
-
How can objects help action recognition? 20 Jun 2023 · 1 repository · arXiv:2306.11726
-
MuDPT: Multi-modal Deep-symphysis Prompt Tuning for Large Pre-trained Vision-Language Models 20 Jun 2023 · 1 repository · arXiv:2306.11400
-
Quilt-1M: One Million Image-Text Pairs for Histopathology 20 Jun 2023 · 2 repositories · arXiv:2306.11207Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
RS5M and GeoRSCLIP: A Large Scale Vision-Language Dataset and A Large Vision-Language Model for Remote Sensing 20 Jun 2023 · 1 repository · arXiv:2306.11300Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 7 harvested samples)
-
RemoteCLIP: A Vision Language Foundation Model for Remote Sensing 19 Jun 2023 · 1 repository · arXiv:2306.11029
-
Instant Soup: Cheap Pruning Ensembles in A Single Pass Can Draw Lottery Tickets from Large Models 18 Jun 2023 · 1 repository · arXiv:2306.10460Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
CLIPSonic: Text-to-Audio Synthesis with Unlabeled Videos and Pretrained Language-Vision Models 16 Jun 2023 · 0 repositories · arXiv:2306.09635
-
Crowdsourcing and Evaluating Text-Based Audio Retrieval Relevances 16 Jun 2023 · 1 repository · arXiv:2306.09820
-
Vision-Language Models can Identify Distracted Driver Behavior from Naturalistic Videos 16 Jun 2023 · 1 repository · arXiv:2306.10159
-
Contrasting Intra-Modal and Ranking Cross-Modal Hard Negatives to Enhance Visio-Linguistic Compositional Understanding 15 Jun 2023 · 2 repositories · arXiv:2306.08832Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Evaluating Data Attribution for Text-to-Image Models 15 Jun 2023 · 2 repositories · arXiv:2306.09345Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Exploring the Application of Large-scale Pre-trained Models on Adverse Weather Removal 15 Jun 2023 · 0 repositories · arXiv:2306.09008
-
Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis 15 Jun 2023 · 1 repository · arXiv:2306.09341Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Pragmatic Inference with a CLIP Listener for Contrastive Captioning 15 Jun 2023 · 1 repository · arXiv:2306.08818
-
Semantic HELM: A Human-Readable Memory for Reinforcement Learning 15 Jun 2023 · 1 repository · arXiv:2306.09312
-
Babel-ImageNet: Massively Multilingual Evaluation of Vision-and-Language Representations 14 Jun 2023 · 1 repository · arXiv:2306.08658Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples)
-
CLIPXPlore: Coupled CLIP and Shape Spaces for 3D Shape Exploration 14 Jun 2023 · 0 repositories · arXiv:2306.08226
-
Extending CLIP's Image-Text Alignment to Referring Image Segmentation 14 Jun 2023 · 0 repositories · arXiv:2306.08498
-
Towards trustworthy seizure onset detection using workflow notes 14 Jun 2023 · 1 repository · arXiv:2306.08728
-
GeneCIS: A Benchmark for General Conditional Image Similarity 13 Jun 2023 · 0 repositories · arXiv:2306.07969
-
Marking anything: application of point cloud in extracting video target features 13 Jun 2023 · 0 repositories · arXiv:2306.07559
-
MOFI: Learning Image Representations from Noisy Entity Annotated Images 13 Jun 2023 · 1 repository · arXiv:2306.07952Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Safeguarding Data in Multimodal AI: A Differentially Private Approach to CLIP Training 13 Jun 2023 · 1 repository · arXiv:2306.08173
-
A Survey of Vision-Language Pre-training from the Lens of Multimodal Machine Translation 12 Jun 2023 · 0 repositories · arXiv:2306.07198
-
Augmenting Zero-Shot Detection Training with Image Labels 12 Jun 2023 · 0 repositories · arXiv:2306.06899
-
Learning to Mask and Permute Visual Tokens for Vision Transformer Pre-Training 12 Jun 2023 · 1 repository · arXiv:2306.07346
-
Retrieval-Enhanced Contrastive Vision-Text Models 12 Jun 2023 · 0 repositories · arXiv:2306.07196
-
Sticker820K: Empowering Interactive Retrieval with Stickers 12 Jun 2023 · 0 repositories · arXiv:2306.06870
-
Waffling around for Performance: Visual Classification with Random Words and Broad Concepts 12 Jun 2023 · 2 repositories · arXiv:2306.07282Syntology official (archive's flag): 4 ran · 5 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 5 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples) · 1 pointer-only (licence)
-
EventCLIP: Adapting CLIP for Event-based Object Recognition 10 Jun 2023 · 1 repository · arXiv:2306.06354Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 10 harvested samples)
-
How Does Fine-Tuning Impact Out-of-Distribution Detection for Vision-Language Models? 9 Jun 2023 · 0 repositories · arXiv:2306.06048
-
Assessing Phrase Break of ESL Speech with Pre-trained Language Models and Large Language Models 8 Jun 2023 · 0 repositories · arXiv:2306.04980
-
Image Clustering via the Principle of Rate Reduction in the Age of Pretrained Models 8 Jun 2023 · 1 repository · arXiv:2306.05272
-
SyncDiffusion: Coherent Montage via Synchronized Joint Diffusions 8 Jun 2023 · 0 repositories · arXiv:2306.05178
-
Fine-Grained Visual Prompting 7 Jun 2023 · 1 repository · arXiv:2306.04356
-
UniBoost: Unsupervised Unimodal Pre-training for Boosting Zero-shot Vision-Language Tasks 7 Jun 2023 · 0 repositories · arXiv:2306.04715
-
Emotional Talking Head Generation based on Memory-Sharing and Attention-Augmented Networks 6 Jun 2023 · 0 repositories · arXiv:2306.03594
-
Identifying Shared Decodable Concepts in the Human Brain Using Image-Language Foundation Models 6 Jun 2023 · 0 repositories · arXiv:2306.03375
-
On the Difference of BERT-style and CLIP-style Text Encoders 6 Jun 2023 · 1 repository · arXiv:2306.03678
-
Recognize Anything: A Strong Image Tagging Model 6 Jun 2023 · 2 repositories · arXiv:2306.03514
-
Towards Label-free Scene Understanding by Vision Foundation Models 6 Jun 2023 · 1 repository · arXiv:2306.03899
-
Semantically-Prompted Language Models Improve Visual Descriptions 5 Jun 2023 · 0 repositories · arXiv:2306.06077
-
Detector Guidance for Multi-Object Text-to-Image Generation 4 Jun 2023 · 1 repository · arXiv:2306.02236
-
MoviePuzzle: Visual Narrative Reasoning through Multimodal Order Learning 4 Jun 2023 · 0 repositories · arXiv:2306.02252
-
Multi-CLIP: Contrastive Vision-Language Pre-training for Question Answering tasks in 3D Scenes 4 Jun 2023 · 0 repositories · arXiv:2306.02329
-
ProTeCt: Prompt Tuning for Taxonomic Open Set Classification 4 Jun 2023 · 1 repository · arXiv:2306.02240Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 1 where Syntology's instrument failed) · 4 unverified (of 10 harvested samples) · 1 pointer-only (licence)