Methods › Computer Vision › Vision and Language Pre-Trained Models › CLIP › Papers, page 6
Contrastive Language-Image Pre-training
CLIP
Papers archive 2025-07-28
archive papers tagged: 3,094 · with a code link: 1,617 · where Syntology ran a sample: 649 (554 with a run with no instrument failure, 95 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (649 of 3,094 tagged: 554 with a run with no instrument failure, 95 where every run was a failure of Syntology's instrument)
Page 6 of 31: papers 501 to 600 of 3,094, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
V2C-CBM: Building Concept Bottlenecks with Vision-to-Concept Tokenizer 9 Jan 2025 · 1 repository · arXiv:2501.04975
-
Vision-Language Models for Autonomous Driving: CLIP-Based Dynamic Scene Understanding 9 Jan 2025 · 0 repositories · arXiv:2501.05566
-
ContextMRI: Enhancing Compressed Sensing MRI through Metadata Conditioning 8 Jan 2025 · 1 repository · arXiv:2501.04284
-
Rethinking High-speed Image Reconstruction Framework with Spike Camera 8 Jan 2025 · 1 repository · arXiv:2501.04477
-
Unified Coding for Both Human Perception and Generalized Machine Analytics with CLIP Supervision 8 Jan 2025 · 1 repository · arXiv:2501.04579
-
Graph-Based Multimodal and Multi-view Alignment for Keystep Recognition 7 Jan 2025 · 1 repository · arXiv:2501.04121
-
KAnoCLIP: Zero-Shot Anomaly Detection through Knowledge-Driven Prompt Learning and Enhanced Cross-Modal Integration 7 Jan 2025 · 0 repositories · arXiv:2501.03786
-
MADation: Face Morphing Attack Detection with Foundation Models 7 Jan 2025 · 2 repositories · arXiv:2501.03800
-
MedFocusCLIP : Improving few shot classification in medical datasets using pixel wise attention 7 Jan 2025 · 0 repositories · arXiv:2501.03839
-
MedicalNarratives: Connecting Medical Vision and Language with Localized Narratives 7 Jan 2025 · 0 repositories · arXiv:2501.04184
-
RAG-Check: Evaluating Multimodal Retrieval Augmented Generation Performance 7 Jan 2025 · 0 repositories · arXiv:2501.03995
-
Self-adaptive vision-language model for 3D segmentation of pulmonary artery and vein 7 Jan 2025 · 0 repositories · arXiv:2501.03722
-
Multilevel Semantic-Aware Model for AI-Generated Video Quality Assessment 6 Jan 2025 · 0 repositories · arXiv:2501.02706
-
Automated Detection of Epileptic Spikes and Seizures Incorporating a Novel Spatial Clustering Prior 5 Jan 2025 · 0 repositories · arXiv:2501.10404
-
Decoding fMRI Data into Captions using Prefix Language Modeling 5 Jan 2025 · 1 repository · arXiv:2501.02570
-
FedRSClip: Federated Learning for Remote Sensing Scene Classification Using Vision-Language Models 5 Jan 2025 · 0 repositories · arXiv:2501.02461
-
A Survey of State of the Art Large Vision Language Models: Alignment, Benchmark, Evaluations and Challenges 4 Jan 2025 · 3 repositories · arXiv:2501.02189
-
CLIP-UP: CLIP-Based Unanswerable Problem Detection for Visual Question Answering 2 Jan 2025 · 0 repositories · arXiv:2501.01371
-
TexAVi: Generating Stereoscopic VR Video Clips from Text Descriptions 2 Jan 2025 · 0 repositories · arXiv:2501.01156
-
A3: Few-shot Prompt Learning of Unlearnable Examples with Cross-Modal Adversarial Feature Alignment 1 Jan 2025 · 0 repositories
-
Adaptive Parameter Selection for Tuning Vision-Language Models 1 Jan 2025 · 0 repositories
-
Advancing Myopia To Holism: Fully Contrastive Language-Image Pre-training 1 Jan 2025 · 0 repositories
-
Bringing CLIP to the Clinic: Dynamic Soft Labels and Negation-Aware Learning for Medical Analysis 1 Jan 2025 · 0 repositories
-
CacheQuant: Comprehensively Accelerated Diffusion Models 1 Jan 2025 · 0 repositories
-
Classifier-guided CLIP Distillation for Unsupervised Multi-label Classification 1 Jan 2025 · 1 repository
-
CLIP-driven Coarse-to-fine Semantic Guidance for Fine-grained Open-set Semi-supervised Learning 1 Jan 2025 · 0 repositories
-
CLIP is Almost All You Need: Towards Parameter-Efficient Scene Text Retrieval without OCR 1 Jan 2025 · 0 repositories
-
Diffusion Bridge: Leveraging Diffusion Model to Reduce the Modality Gap Between Text and Vision for Zero-Shot Image Captioning 1 Jan 2025 · 1 repository
-
Domain Generalization in CLIP via Learning with Diverse Text Prompts 1 Jan 2025 · 0 repositories
-
DTOS: Dynamic Time Object Sensing with Large Multimodal Model 1 Jan 2025 · 1 repository
-
Dual Semantic Guidance for Open Vocabulary Semantic Segmentation 1 Jan 2025 · 0 repositories
-
Enhancing Diversity for Data-free Quantization 1 Jan 2025 · 0 repositories
-
Enhancing Few-Shot Class-Incremental Learning via Training-Free Bi-Level Modality Calibration 1 Jan 2025 · 1 repository
-
FGAseg: Fine-Grained Pixel-Text Alignment for Open-Vocabulary Semantic Segmentation 1 Jan 2025 · 1 repository · arXiv:2501.00877
-
Forensics Adapter: Adapting CLIP for Generalizable Face Forgery Detection 1 Jan 2025 · 0 repositories
-
GET: Unlocking the Multi-modal Potential of CLIP for Generalized Category Discovery 1 Jan 2025 · 0 repositories
-
Harnessing Frozen Unimodal Encoders for Flexible Multimodal Alignment 1 Jan 2025 · 1 repository
-
Hierarchical Knowledge Prompt Tuning for Multi-task Test-Time Adaptation 1 Jan 2025 · 0 repositories
-
HORUS: Multimodal Large Language Models Framework for Video Retrieval at VBS 2025 1 Jan 2025 · 0 repositories
-
ImagineFSL: Self-Supervised Pretraining Matters on Imagined Base Set for VLM-based Few-shot Learning 1 Jan 2025 · 1 repository
-
Incorporating Dense Knowledge Alignment into Unified Multimodal Representation Models 1 Jan 2025 · 0 repositories
-
LOGICZSL: Exploring Logic-induced Representation for Compositional Zero-shot Learning 1 Jan 2025 · 0 repositories
-
On the Zero-shot Adversarial Robustness of Vision-Language Models: A Truly Zero-shot and Training-free Approach 1 Jan 2025 · 0 repositories
-
Open Ad-hoc Categorization with Contextualized Feature Learning 1 Jan 2025 · 0 repositories
-
Overcoming Shortcut Problem in VLM for Robust Out-of-Distribution Detection 1 Jan 2025 · 1 repository
-
Preserving Clusters in Prompt Learning for Unsupervised Domain Adaptation 1 Jan 2025 · 0 repositories
-
RADIOv2.5: Improved Baselines for Agglomerative Vision Foundation Models 1 Jan 2025 · 0 repositories
-
Retaining Knowledge and Enhancing Long-Text Representations in CLIP through Dual-Teacher Distillation 1 Jan 2025 · 0 repositories
-
SmartCLIP: Modular Vision-language Alignment with Identification Guarantees 1 Jan 2025 · 1 repository
-
SpatialCLIP: Learning 3D-aware Image Representations from Spatially Discriminative Language 1 Jan 2025 · 0 repositories
-
Style-Editor: Text-driven Object-centric Style Editing 1 Jan 2025 · 0 repositories
-
T2ICount: Enhancing Cross-modal Understanding for Zero-Shot Counting 1 Jan 2025 · 1 repository
-
Targeted Forgetting of Image Subgroups in CLIP Models 1 Jan 2025 · 0 repositories
-
Text Augmented Correlation Transformer For Few-shot Classification & Segmentation 1 Jan 2025 · 0 repositories
-
Towards More General Video-based Deepfake Detection through Facial Component Guided Adaptation for Foundation Model 1 Jan 2025 · 1 repository
-
Understanding Fine-tuning CLIP for Open-vocabulary Semantic Segmentation in Hyperbolic Space 1 Jan 2025 · 0 repositories
-
Differentiable Prompt Learning for Vision Language Models 31 Dec 2024 · 0 repositories · arXiv:2501.00457
-
Dynamic Prompt Adjustment for Multi-Label Class-Incremental Learning 31 Dec 2024 · 0 repositories · arXiv:2501.00340
-
Image Fusion for Cross-Domain Sequential Recommendation 31 Dec 2024 · 0 repositories · arXiv:2502.15694
-
Unleashing Text-to-Image Diffusion Prior for Zero-Shot Image Captioning 31 Dec 2024 · 0 repositories · arXiv:2501.00437
-
AltGen: AI-Driven Alt Text Generation for Enhancing EPUB Accessibility 30 Dec 2024 · 0 repositories · arXiv:2501.00113
-
E2EDiff: Direct Mapping from Noise to Data for Enhanced Diffusion Models 30 Dec 2024 · 0 repositories · arXiv:2412.21044
-
Enhancing Visual Representation for Text-based Person Searching 30 Dec 2024 · 1 repository · arXiv:2412.20646
-
Learning to Rank Pre-trained Vision-Language Models for Downstream Tasks 30 Dec 2024 · 0 repositories · arXiv:2412.20682
-
Towards Compatible Fine-tuning for Vision-Language Model Updates 30 Dec 2024 · 0 repositories · arXiv:2412.20895
-
Towards Identity-Aware Cross-Modal Retrieval: a Dataset and a Baseline 30 Dec 2024 · 1 repository · arXiv:2412.21009
-
YOLO-UniOW: Efficient Universal Open-World Object Detection 30 Dec 2024 · 1 repository · arXiv:2412.20645
-
Defending Multimodal Backdoored Models by Repulsive Visual Prompt Tuning 29 Dec 2024 · 0 repositories · arXiv:2412.20392
-
Cross-Modal Mapping: Mitigating the Modality Gap for Few-Shot Image Classification 28 Dec 2024 · 0 repositories · arXiv:2412.20110
-
Injecting Explainability and Lightweight Design into Weakly Supervised Video Anomaly Detection Systems 28 Dec 2024 · 0 repositories · arXiv:2412.20201
-
Multi-Modality Driven LoRA for Adverse Condition Depth Estimation 28 Dec 2024 · 0 repositories · arXiv:2412.20162
-
Data-Free Group-Wise Fully Quantized Winograd Convolution via Learnable Scales 27 Dec 2024 · 0 repositories · arXiv:2412.19867
-
ReNeg: Learning Negative Embedding with Reward Guidance 27 Dec 2024 · 2 repositories · arXiv:2412.19637
-
Toward Modality Gap: Vision Prototype Learning for Weakly-supervised Semantic Segmentation with CLIP 27 Dec 2024 · 0 repositories · arXiv:2412.19650
-
CLIP-GS: Unifying Vision-Language Representation with 3D Gaussian Splatting 26 Dec 2024 · 0 repositories · arXiv:2412.19142
-
From Coin to Data: The Impact of Object Detection on Digital Numismatics 26 Dec 2024 · 0 repositories · arXiv:2412.19091
-
Referencing Where to Focus: Improving VisualGrounding with Referential Query 26 Dec 2024 · 0 repositories · arXiv:2412.19155
-
ViPCap: Retrieval Text-Based Visual Prompts for Lightweight Image Captioning 26 Dec 2024 · 1 repository · arXiv:2412.19289
-
CLIP-Based Modality Compensation for Visible-Infrared Image Re-Identification 25 Dec 2024 · 0 repositories
-
FOR: Finetuning for Object Level Open Vocabulary Image Retrieval 25 Dec 2024 · 0 repositories · arXiv:2412.18806
-
Open-Vocabulary Panoptic Segmentation Using BERT Pre-Training of Vision-Language Multiway Transformer Model 25 Dec 2024 · 1 repository · arXiv:2412.18917
-
Dissecting CLIP: Decomposition with a Schur Complement-based Approach 24 Dec 2024 · 1 repository · arXiv:2412.18645
-
Extract Free Dense Misalignment from CLIP 24 Dec 2024 · 1 repository · arXiv:2412.18404
-
Sampling Bag of Views for Open-Vocabulary Object Detection 24 Dec 2024 · 0 repositories · arXiv:2412.18273
-
AFANet: Adaptive Frequency-Aware Network for Weakly-Supervised Few-Shot Semantic Segmentation 23 Dec 2024 · 1 repository · arXiv:2412.17601
-
Multimodal Preference Data Synthetic Alignment with Reward Model 23 Dec 2024 · 1 repository · arXiv:2412.17417
-
Bridging Auditory Perception and Language Comprehension through MEG-Driven Encoding Models 22 Dec 2024 · 0 repositories · arXiv:2501.03246
-
MVREC: A General Few-shot Defect Classification Model Using Multi-View Region-Context 22 Dec 2024 · 1 repository · arXiv:2412.16897Syntology official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 11 with no instrument failure: 0 honoured, 0 violated, 11 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 12 harvested samples)
-
SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization 21 Dec 2024 · 1 repository · arXiv:2412.16771
-
A New Method to Capturing Compositional Knowledge in Linguistic Space 20 Dec 2024 · 0 repositories · arXiv:2412.15632
-
Diffusion-Based Conditional Image Editing through Optimized Inference with Guidance 20 Dec 2024 · 0 repositories · arXiv:2412.15798
-
DINOv2 Meets Text: A Unified Framework for Image- and Pixel-Level Vision-Language Alignment 20 Dec 2024 · 1 repository · arXiv:2412.16334
-
SGTC: Semantic-Guided Triplet Co-training for Sparsely Annotated Semi-Supervised Medical Image Segmentation 20 Dec 2024 · 1 repository · arXiv:2412.15526
-
DiffSim: Taming Diffusion Models for Evaluating Visual Similarity 19 Dec 2024 · 1 repository · arXiv:2412.14580
-
Learning Visual Composition through Improved Semantic Guidance 19 Dec 2024 · 0 repositories · arXiv:2412.15396
-
MitraClip Device Automated Localization in 3D Transesophageal Echocardiography via Deep Learning 19 Dec 2024 · 0 repositories · arXiv:2412.15013
-
Multimodal Hypothetical Summary for Retrieval-based Multi-image Question Answering 19 Dec 2024 · 1 repository · arXiv:2412.14880
-
Relational Programming with Foundation Models 19 Dec 2024 · 0 repositories · arXiv:2412.14515
-
GAGS: Granularity-Aware Feature Distillation for Language Gaussian Splatting 18 Dec 2024 · 0 repositories · arXiv:2412.13654
-
I0T: Embedding Standardization Method Towards Zero Modality Gap 18 Dec 2024 · 1 repository · arXiv:2412.14384