Methods › Computer Vision › Vision and Language Pre-Trained Models › CLIP › Papers, page 7
Contrastive Language-Image Pre-training
CLIP
Papers archive 2025-07-28
archive papers tagged: 3,094 · with a code link: 1,617 · where Syntology ran a sample: 649 (554 with a run with no instrument failure, 95 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (649 of 3,094 tagged: 554 with a run with no instrument failure, 95 where every run was a failure of Syntology's instrument)
Page 7 of 31: papers 601 to 700 of 3,094, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
MATCHED: Multimodal Authorship-Attribution To Combat Human Trafficking in Escort-Advertisement Data 18 Dec 2024 · 1 repository · arXiv:2412.13794
-
PLPP: Prompt Learning with Perplexity Is Self-Distillation for Vision-Language Models 18 Dec 2024 · 0 repositories · arXiv:2412.15277
-
Real Classification by Description: Extending CLIP's Limits of Part Attributes Recognition 18 Dec 2024 · 0 repositories · arXiv:2412.13947
-
Beyond Accuracy: On the Effects of Fine-tuning Towards Vision-Language Model's Prediction Rationality 17 Dec 2024 · 1 repository · arXiv:2412.13333
-
CLIP-RLDrive: Human-Aligned Autonomous Driving via CLIP-Based Reward Shaping in Reinforcement Learning 17 Dec 2024 · 0 repositories · arXiv:2412.16201
-
CRoF: CLIP-based Robust Few-shot Learning on Noisy Labels 17 Dec 2024 · 0 repositories · arXiv:2412.12793
-
Optimized two-stage AI-based Neural Decoding for Enhanced Visual Stimulus Reconstruction from fMRI Data 17 Dec 2024 · 0 repositories · arXiv:2412.13237
-
A LoRA is Worth a Thousand Pictures 16 Dec 2024 · 0 repositories · arXiv:2412.12048
-
CLIP-SR: Collaborative Linguistic and Image Processing for Super-Resolution 16 Dec 2024 · 0 repositories · arXiv:2412.11609
-
CPath-Omni: A Unified Multimodal Foundation Model for Patch and Whole Slide Image Analysis in Computational Pathology 16 Dec 2024 · 0 repositories · arXiv:2412.12077
-
Does it Chug? Towards a Data-Driven Understanding of Guitar Tone Description 16 Dec 2024 · 1 repository · arXiv:2412.11769
-
Does VLM Classification Benefit from LLM Description Semantics? 16 Dec 2024 · 1 repository · arXiv:2412.11917Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
LMM-Regularized CLIP Embeddings for Image Classification 16 Dec 2024 · 0 repositories · arXiv:2412.11663
-
MaskCLIP++: A Mask-Based CLIP Fine-tuning Framework for Open-Vocabulary Image Segmentation 16 Dec 2024 · 1 repository · arXiv:2412.11464
-
Text and Image Are Mutually Beneficial: Enhancing Training-Free Few-Shot Classification with CLIP 16 Dec 2024 · 1 repository · arXiv:2412.11375
-
Enhance Vision-Language Alignment with Noise 14 Dec 2024 · 1 repository · arXiv:2412.10817Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 2 unverified (of 5 harvested samples) · 3 pointer-only (licence)
-
MambaPro: Multi-Modal Object Re-Identification with Mamba Aggregation and Synergistic Prompt 14 Dec 2024 · 1 repository · arXiv:2412.10707
-
Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP 13 Dec 2024 · 1 repository · arXiv:2412.09895
-
CognitionCapturer: Decoding Visual Stimuli From Human EEG Signal With Multimodal Information 13 Dec 2024 · 1 repository · arXiv:2412.10489
-
Prompt-Guided Mask Proposal for Two-Stage Open-Vocabulary Segmentation 13 Dec 2024 · 0 repositories · arXiv:2412.10292
-
Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics 13 Dec 2024 · 1 repository · arXiv:2412.10594
-
Agent-based Video Trimming 12 Dec 2024 · 0 repositories · arXiv:2412.09513
-
BayesAdapter: enhanced uncertainty estimation in CLIP few-shot adaptation 12 Dec 2024 · 0 repositories · arXiv:2412.09718
-
Embeddings are all you need! Achieving High Performance Medical Image Classification through Training-Free Embedding Analysis 12 Dec 2024 · 0 repositories · arXiv:2412.09445
-
Omni-ID: Holistic Identity Representation Designed for Generative Tasks 12 Dec 2024 · 0 repositories · arXiv:2412.09694
-
Can Graph Neural Networks Learn Language with Extremely Weak Text Supervision? 11 Dec 2024 · 1 repository · arXiv:2412.08174Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 5 harvested samples)
-
EOV-Seg: Efficient Open-Vocabulary Panoptic Segmentation 11 Dec 2024 · 1 repository · arXiv:2412.08628
-
jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images 11 Dec 2024 · 0 repositories · arXiv:2412.08802
-
POINTS1.5: Building a Vision-Language Model towards Real World Applications 11 Dec 2024 · 0 repositories · arXiv:2412.08443
-
Position-aware Guided Point Cloud Completion with CLIP Model 11 Dec 2024 · 0 repositories · arXiv:2412.08271
-
Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models 11 Dec 2024 · 0 repositories · arXiv:2412.08111
-
SenCLIP: Enhancing zero-shot land-use mapping for Sentinel-2 with ground-level prompting 11 Dec 2024 · 1 repository · arXiv:2412.08536
-
SLGaussian: Fast Language Gaussian Splatting in Sparse Views 11 Dec 2024 · 0 repositories · arXiv:2412.08331
-
AmCLR: Unified Augmented Learning for Cross-Modal Representations 10 Dec 2024 · 1 repository · arXiv:2412.07979
-
Attention Head Purification: A New Perspective to Harness CLIP for Domain Generalization 10 Dec 2024 · 0 repositories · arXiv:2412.07226
-
Cloud Object Detector Adaptation by Integrating Different Source Knowledge 10 Dec 2024 · 1 repository
-
DiffCLIP: Few-shot Language-driven Multimodal Classifier 10 Dec 2024 · 1 repository · arXiv:2412.07119
-
Explaining and Mitigating the Modality Gap in Contrastive Multimodal Learning 10 Dec 2024 · 0 repositories · arXiv:2412.07909
-
Fusion Embedding for Pose-Guided Person Image Synthesis with Diffusion Model 10 Dec 2024 · 0 repositories · arXiv:2412.07333
-
Hero-SR: One-Step Diffusion for Super-Resolution with Human Perception Priors 10 Dec 2024 · 0 repositories · arXiv:2412.07152
-
Leveraging Content and Context Cues for Low-Light Image Enhancement 10 Dec 2024 · 1 repository · arXiv:2412.07693
-
Mobile Video Diffusion 10 Dec 2024 · 0 repositories · arXiv:2412.07583
-
Multimodal Contextualized Support for Enhancing Video Retrieval System 10 Dec 2024 · 0 repositories · arXiv:2412.07584
-
RADIO Amplified: Improved Baselines for Agglomerative Vision Foundation Models 10 Dec 2024 · 1 repository · arXiv:2412.07679
-
Retaining and Enhancing Pre-trained Knowledge in Vision-Language Models with Prompt Ensembling 10 Dec 2024 · 0 repositories · arXiv:2412.07077
-
Category-Adaptive Cross-Modal Semantic Refinement and Transfer for Open-Vocabulary Multi-Label Recognition 9 Dec 2024 · 0 repositories · arXiv:2412.06190
-
DenseVLM: A Retrieval and Decoupled Alignment Framework for Open-Vocabulary Dense Prediction 9 Dec 2024 · 0 repositories · arXiv:2412.06244
-
Ranking-aware adapter for text-driven image ordering with CLIP 9 Dec 2024 · 1 repository · arXiv:2412.06760
-
ZeroKey: Point-Level Reasoning and Zero-Shot 3D Keypoint Detection from Large Language Models 9 Dec 2024 · 0 repositories · arXiv:2412.06292
-
LVP-CLIP:Revisiting CLIP for Continual Learning with Label Vector Pool 8 Dec 2024 · 0 repositories · arXiv:2412.05840
-
Post-hoc Probabilistic Vision-Language Models 8 Dec 2024 · 1 repository · arXiv:2412.06014Syntology official (archive's flag): 4 ran · 7 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 2 honoured, 0 violated, 0 with no contract checked; 5 where Syntology's instrument failed) · 6 unverified (of 13 harvested samples) · 5 pointer-only (licence)
-
CLIP-TNseg: A Multi-Modal Hybrid Framework for Thyroid Nodule Segmentation in Ultrasound Images 7 Dec 2024 · 1 repository · arXiv:2412.05530
-
Compositional Image Retrieval via Instruction-Aware Contrastive Learning 7 Dec 2024 · 1 repository · arXiv:2412.05756
-
Parametric-ControlNet: Multimodal Control in Foundation Models for Precise Engineering Design Synthesis 6 Dec 2024 · 0 repositories · arXiv:2412.04707
-
S³: Synonymous Semantic Space for Improving Zero-Shot Generalization of Vision-Language Models 6 Dec 2024 · 0 repositories · arXiv:2412.04925
-
SMIC: Semantic Multi-Item Compression based on CLIP dictionary 6 Dec 2024 · 0 repositories · arXiv:2412.05035
-
Sparse autoencoders reveal selective remapping of visual concepts during adaptation 6 Dec 2024 · 1 repository · arXiv:2412.05276Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 6 harvested samples)
-
4Real-Video: Learning Generalizable Photo-Realistic 4D Video Diffusion 5 Dec 2024 · 0 repositories · arXiv:2412.04462
-
Assessing and Learning Alignment of Unimodal Vision and Language Models 5 Dec 2024 · 0 repositories · arXiv:2412.04616
-
CLIP-FSAC++: Few-Shot Anomaly Classification with Anomaly Descriptor Based on CLIP 5 Dec 2024 · 0 repositories · arXiv:2412.03829
-
CLIP-PING: Boosting Lightweight Vision-Language Models with Proximus Intrinsic Neighbors Guidance 5 Dec 2024 · 0 repositories · arXiv:2412.03871
-
VladVA: Discriminative Fine-tuning of LVLMs 5 Dec 2024 · 0 repositories · arXiv:2412.04378
-
Grounding Descriptions in Images informs Zero-Shot Visual Recognition 5 Dec 2024 · 1 repository · arXiv:2412.04429
-
Liquid: Language Models are Scalable Multi-modal Generators 5 Dec 2024 · 1 repository · arXiv:2412.04332Syntology official (archive's flag): 1 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 1 pointer-only (licence)
-
Mask-Adapter: The Devil is in the Masks for Open-Vocabulary Segmentation 5 Dec 2024 · 1 repository · arXiv:2412.04533
-
VisionZip: Longer is Better but Not Necessary in Vision Language Models 5 Dec 2024 · 1 repository · arXiv:2412.04467Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 1 where Syntology's instrument failed) · 3 unverified (of 11 harvested samples)
-
FLAIR: VLM with Fine-grained Language-informed Image Representations 4 Dec 2024 · 2 repositories · arXiv:2412.03561Syntology official (archive's flag): 13 ran · 13 ran (of which 0 constructed an object rather than computing a result; 11 with no instrument failure: 0 honoured, 1 violated, 10 with no contract checked; 2 where Syntology's instrument failed) · 7 unverified (of 20 harvested samples) · 20 pointer-only (licence)
-
Enhancing CLIP Conceptual Embedding through Knowledge Distillation 4 Dec 2024 · 0 repositories · arXiv:2412.03513
-
Training-Free Mitigation of Language Reasoning Degradation After Multimodal Instruction Tuning 4 Dec 2024 · 0 repositories · arXiv:2412.03467
-
Enhancing Robustness of CLIP to Common Corruptions through Bimodal Test-Time Adaptation 3 Dec 2024 · 0 repositories · arXiv:2412.02837
-
Gaussian Splatting Under Attack: Investigating Adversarial Noise in 3D Objects 3 Dec 2024 · 1 repository · arXiv:2412.02803Syntology 0 ran · 2 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Improving Dynamic Object Interactions in Text-to-Video Generation with AI Feedback 3 Dec 2024 · 0 repositories · arXiv:2412.02617
-
SparseLGS: Sparse View Language Embedded Gaussian Splatting 3 Dec 2024 · 0 repositories · arXiv:2412.02245
-
Viewpoint Consistency in 3D Generation via Attention and CLIP Guidance 3 Dec 2024 · 0 repositories · arXiv:2412.02287
-
3DSceneEditor: Controllable 3D Scene Editing with Gaussian Splatting 2 Dec 2024 · 0 repositories · arXiv:2412.01583
-
Attacks on multimodal models 2 Dec 2024 · 1 repository · arXiv:2412.01725
-
BroadTrack: Broadcast Camera Tracking for Soccer 2 Dec 2024 · 1 repository · arXiv:2412.01721
-
MuLan: Adapting Multilingual Diffusion Models for Hundreds of Languages with Negligible Cost 2 Dec 2024 · 1 repository · arXiv:2412.01271
-
NLPrompt: Noise-Label Prompt Learning for Vision-Language Models 2 Dec 2024 · 1 repository · arXiv:2412.01256Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 3 where Syntology's instrument failed) · 2 unverified (of 9 harvested samples) · 9 pointer-only (licence)
-
See What You Seek: Semantic Contextual Integration for Cloth-Changing Person Re-Identification 2 Dec 2024 · 0 repositories · arXiv:2412.01345
-
VideoLights: Feature Refinement and Cross-Task Alignment Transformer for Joint Video Highlight Detection and Moment Retrieval 2 Dec 2024 · 1 repository · arXiv:2412.01558
-
Adaptive Rank, Reduced Forgetting: Knowledge Retention in Continual Learning Vision-Language Models with Dynamic Rank-Selective LoRA 1 Dec 2024 · 0 repositories · arXiv:2412.01004
-
Perturb and Recover: Fine-tuning for Effective Backdoor Removal from CLIP 1 Dec 2024 · 1 repository · arXiv:2412.00727
-
Prompt as Free Lunch: Enhancing Diversity in Source-Free Cross-domain Few-shot Learning through Semantic-Guided Prompting 1 Dec 2024 · 0 repositories · arXiv:2412.00767
-
STEVE-Audio: Expanding the Goal Conditioning Modalities of Embodied Agents in Minecraft 1 Dec 2024 · 0 repositories · arXiv:2412.00949
-
LMSeg: Unleashing the Power of Large-Scale Models for Open-Vocabulary Semantic Segmentation 30 Nov 2024 · 0 repositories · arXiv:2412.00364
-
Bootstraping Clustering of Gaussians for View-consistent 3D Scene Understanding 29 Nov 2024 · 1 repository · arXiv:2411.19551Syntology official (archive's flag): 12 ran · 12 ran (of which 0 constructed an object rather than computing a result; 11 with no instrument failure: 2 honoured, 0 violated, 9 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 12 harvested samples) · 12 pointer-only (licence)
-
Dual Risk Minimization: Towards Next-Level Robustness in Fine-tuning Zero-Shot Models 29 Nov 2024 · 1 repository · arXiv:2411.19757Syntology official (archive's flag): 6 ran · 6 ran (of which 1 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 4 where Syntology's instrument failed) · 5 unverified (of 11 harvested samples) · 11 pointer-only (licence)
-
Forensics Adapter: Unleashing CLIP for Generalizable Face Forgery Detection 29 Nov 2024 · 1 repository · arXiv:2411.19715
-
GuardSplat: Efficient and Robust Watermarking for 3D Gaussian Splatting 29 Nov 2024 · 1 repository · arXiv:2411.19895
-
ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model 29 Nov 2024 · 0 repositories · arXiv:2412.00153
-
Automatic Prompt Generation and Grounding Object Detection for Zero-Shot Image Anomaly Detection 28 Nov 2024 · 0 repositories · arXiv:2411.19220
-
CLIP meets DINO for Tuning Zero-Shot Classifier using Unlabeled Image Collections 28 Nov 2024 · 1 repository · arXiv:2411.19346
-
Talking to DINO: Bridging Self-Supervised Vision Backbones with Language for Open-Vocabulary Segmentation 28 Nov 2024 · 1 repository · arXiv:2411.19331Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 10 harvested samples) · 2 pointer-only (licence)
-
GLS: Geometry-aware 3D Language Gaussian Splatting 27 Nov 2024 · 0 repositories · arXiv:2411.18066
-
Reconstructing Animals and the Wild 27 Nov 2024 · 0 repositories · arXiv:2411.18807
-
VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis 27 Nov 2024 · 0 repositories · arXiv:2411.18038
-
FLEX-CLIP: Feature-Level GEneration Network Enhanced CLIP for X-shot Cross-modal Retrieval 26 Nov 2024 · 0 repositories · arXiv:2411.17454
-
Words Matter: Leveraging Individual Text Embeddings for Code Generation in CLIP Test-Time Adaptation 26 Nov 2024 · 1 repository · arXiv:2411.17002
-
CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions 25 Nov 2024 · 0 repositories · arXiv:2411.16828