Methods › Computer Vision › Vision and Language Pre-Trained Models › CLIP › Papers, page 3
Contrastive Language-Image Pre-training
CLIP
Papers archive 2025-07-28
archive papers tagged: 3,094 · with a code link: 1,617 · where Syntology ran a sample: 649 (554 with a run with no instrument failure, 95 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (649 of 3,094 tagged: 554 with a run with no instrument failure, 95 where every run was a failure of Syntology's instrument)
Page 3 of 31: papers 201 to 300 of 3,094, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
The 1st EReL@MIR Workshop on Efficient Representation Learning for Multimodal Information Retrieval 21 Apr 2025 · 0 repositories · arXiv:2504.14788
-
sEEG-based Encoding for Sentence Retrieval: A Contrastive Learning Approach to Brain-Language Alignment 20 Apr 2025 · 0 repositories · arXiv:2504.14468
-
CLIP-Powered Domain Generalization and Domain Adaptation: A Comprehensive Survey 19 Apr 2025 · 1 repository · arXiv:2504.14280
-
Cross-attention for State-based model RWKV-7 19 Apr 2025 · 1 repository · arXiv:2504.14260
-
Revisiting CLIP for SF-OSDA: Unleashing Zero-Shot Potential with Adaptive Threshold and Training-Free Feature Filtering 19 Apr 2025 · 0 repositories · arXiv:2504.14224
-
LoftUp: Learning a Coordinate-Based Feature Upsampler for Vision Foundation Models 18 Apr 2025 · 2 repositories · arXiv:2504.14032Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 2 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 6 harvested samples) · 1 pointer-only (licence)
-
ProgRoCC: A Progressive Approach to Rough Crowd Counting 18 Apr 2025 · 0 repositories · arXiv:2504.13405
-
Towards a Multi-Agent Vision-Language System for Zero-Shot Novel Hazardous Object Detection for Autonomous Driving Safety 18 Apr 2025 · 1 repository · arXiv:2504.13399Syntology official: harvested, nothing ran · 0 ran · 2 unverified (of 2 harvested samples)
-
Post-pre-training for Modality Alignment in Vision-Language Foundation Models 17 Apr 2025 · 1 repository · arXiv:2504.12717
-
Science-T2I: Addressing Scientific Illusions in Image Synthesis 17 Apr 2025 · 0 repositories · arXiv:2504.13129
-
AdaVid: Adaptive Video-Language Pretraining 16 Apr 2025 · 0 repositories · arXiv:2504.12513
-
DVLTA-VQA: Decoupled Vision-Language Modeling with Text-Guided Adaptation for Blind Video Quality Assessment 16 Apr 2025 · 0 repositories · arXiv:2504.11733
-
Logits DeConfusion with CLIP for Few-Shot Learning 16 Apr 2025 · 1 repository · arXiv:2504.12104Syntology official (archive's flag): 6 ran · 6 ran (of which 6 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified; every one of the 6 samples that ran constructed an object rather than computing a result (of 7 harvested samples) · 7 pointer-only (licence)
-
Co-STAR: Collaborative Curriculum Self-Training with Adaptive Regularization for Source-Free Video Domain Adaptation 15 Apr 2025 · 0 repositories · arXiv:2504.11669
-
Crane: Context-Guided Prompt Learning and Attention Refinement for Zero-Shot Anomaly Detections 15 Apr 2025 · 1 repository · arXiv:2504.11055
-
R-TPT: Improving Adversarial Robustness of Vision-Language Models through Test-Time Prompt Tuning 15 Apr 2025 · 1 repository · arXiv:2504.11195Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
TMCIR: Token Merge Benefits Composed Image Retrieval 15 Apr 2025 · 0 repositories · arXiv:2504.10995
-
Towards Efficient Partially Relevant Video Retrieval with Active Moment Discovering 15 Apr 2025 · 1 repository · arXiv:2504.10920Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
FLOSS: Free Lunch in Open-vocabulary Semantic Segmentation 14 Apr 2025 · 1 repository · arXiv:2504.10487
-
UP-Person: Unified Parameter-Efficient Transfer Learning for Text-based Person Retrieval 14 Apr 2025 · 1 repository · arXiv:2504.10084
-
3D CoCa: Contrastive Learners are 3D Captioners 13 Apr 2025 · 1 repository · arXiv:2504.09518
-
AeroLite: Tag-Guided Lightweight Generation of Aerial Image Captions 13 Apr 2025 · 0 repositories · arXiv:2504.09528
-
Automatic Detection of Intro and Credits in Video using CLIP and Multihead Attention 13 Apr 2025 · 0 repositories · arXiv:2504.09738
-
Probability Distribution Alignment and Low-Rank Weight Decomposition for Source-Free Domain Adaptive Brain Decoding 12 Apr 2025 · 0 repositories · arXiv:2504.09109
-
FocalLens: Instruction Tuning Enables Zero-Shot Conditional Image Representations 11 Apr 2025 · 0 repositories · arXiv:2504.08368
-
HyperCore: The Core Framework for Building Hyperbolic Foundation Models with Comprehensive Modules 11 Apr 2025 · 1 repository · arXiv:2504.08912Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
self-prompting analogical reasoning for uav object detection 11 Apr 2025 · 1 repository
-
VL-UR: Vision-Language-guided Universal Restoration of Images Degraded by Adverse Weather Conditions 11 Apr 2025 · 0 repositories · arXiv:2504.08219
-
FMNV: A Dataset of Media-Published News Videos for Fake News Detection 10 Apr 2025 · 0 repositories · arXiv:2504.07687
-
Gen3DEval: Using vLLMs for Automatic Evaluation of Generated 3D Objects 10 Apr 2025 · 0 repositories · arXiv:2504.08125
-
Impact of Language Guidance: A Reproducibility Study 10 Apr 2025 · 0 repositories · arXiv:2504.08140
-
MultiADS: Defect-aware Supervision for Multi-type Anomaly Detection and Segmentation in Zero-Shot Learning 9 Apr 2025 · 0 repositories · arXiv:2504.06740
-
Text-to-Image Models and Their Representation of People from Different Nationalities Engaging in Activities 8 Apr 2025 · 0 repositories · arXiv:2504.06313
-
econSG: Efficient and Multi-view Consistent Open-Vocabulary 3D Semantic Gaussians 8 Apr 2025 · 0 repositories · arXiv:2504.06003
-
DA2Diff: Exploring Degradation-aware Adaptive Diffusion Priors for All-in-One Weather Restoration 7 Apr 2025 · 0 repositories · arXiv:2504.05135
-
SDAFE: A Dual-filter Stable Diffusion Data Augmentation Method for Facial Expression Recognition 6 Apr 2025 · 0 repositories
-
Learning Sparse Disentangled Representations for Multimodal Exclusion Retrieval 4 Apr 2025 · 0 repositories · arXiv:2504.03184
-
AC-LoRA: Auto Component LoRA for Personalized Artistic Style Image Generation 3 Apr 2025 · 0 repositories · arXiv:2504.02231
-
Refining CLIP's Spatial Awareness: A Visual-Centric Perspective 3 Apr 2025 · 0 repositories · arXiv:2504.02328
-
Sparse Autoencoders Learn Monosemantic Features in Vision-Language Models 3 Apr 2025 · 1 repository · arXiv:2504.02821Syntology official (archive's flag): 8 ran · 8 ran (of which 3 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 5 where Syntology's instrument failed) · 2 unverified (of 10 harvested samples) · 10 pointer-only (licence)
-
AdPO: Enhancing the Adversarial Robustness of Large Vision-Language Models with Preference Optimization 2 Apr 2025 · 0 repositories · arXiv:2504.01735
-
CLIP-SLA: Parameter-Efficient CLIP Adaptation for Continuous Sign Language Recognition 2 Apr 2025 · 1 repository · arXiv:2504.01666
-
DALIP: Distribution Alignment-based Language-Image Pre-Training for Domain-Specific Data 2 Apr 2025 · 0 repositories · arXiv:2504.01386
-
FineLIP: Extending CLIP's Reach via Fine-Grained Alignment with Longer Text Inputs 2 Apr 2025 · 0 repositories · arXiv:2504.01916
-
Is Temporal Prompting All We Need For Limited Labeled Action Recognition? 2 Apr 2025 · 0 repositories · arXiv:2504.01890
-
Hybrid Global-Local Representation with Augmented Spatial Guidance for Zero-Shot Referring Image Segmentation 1 Apr 2025 · 1 repository · arXiv:2504.00356
-
Knowledge-Base based Semantic Image Transmission Using CLIP 1 Apr 2025 · 0 repositories · arXiv:2504.01053
-
SMILE: Infusing Spatial and Motion Semantics in Masked Video Learning 1 Apr 2025 · 1 repository · arXiv:2504.00527
-
Unleashing the Power of Pre-trained Encoders for Universal Adversarial Attack Detection 1 Apr 2025 · 0 repositories · arXiv:2504.00429
-
Zero-Shot 4D Lidar Panoptic Segmentation 1 Apr 2025 · 0 repositories · arXiv:2504.00848
-
CIBR: Cross-modal Information Bottleneck Regularization for Robust CLIP Generalization 31 Mar 2025 · 0 repositories · arXiv:2503.24182
-
Crossmodal Knowledge Distillation with WordNet-Relaxed Text Embeddings for Robust Image Classification 31 Mar 2025 · 0 repositories · arXiv:2503.24017
-
LATex: Leveraging Attribute-based Text Knowledge for Aerial-Ground Person Re-Identification 31 Mar 2025 · 0 repositories · arXiv:2503.23722
-
Order Matters: On Parameter-Efficient Image-to-Video Probing for Recognizing Nearly Symmetric Actions 31 Mar 2025 · 0 repositories · arXiv:2503.24298
-
The Devil is in the Distributions: Explicit Modeling of Scene Content is Key in Zero-Shot Video Captioning 31 Mar 2025 · 0 repositories · arXiv:2503.23679
-
COSMIC: Clique-Oriented Semantic Multi-space Integration for Robust CLIP Test-Time Adaptation 30 Mar 2025 · 1 repository · arXiv:2503.23388Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample)
-
Embedding Shift Dissection on CLIP: Effects of Augmentations on VLM's Representation Learning 30 Mar 2025 · 0 repositories · arXiv:2503.23495
-
Language Guided Concept Bottleneck Models for Interpretable Continual Learning 30 Mar 2025 · 1 repository · arXiv:2503.23283
-
ReasonGrounder: LVLM-Guided Hierarchical Feature Splatting for Open-Vocabulary 3D Visual Grounding and Reasoning 30 Mar 2025 · 0 repositories · arXiv:2503.23297
-
Semantic-Spatial Feature Fusion with Dynamic Graph Refinement for Remote Sensing Image Captioning 30 Mar 2025 · 0 repositories · arXiv:2503.23453
-
Agent-Centric Personalized Multiple Clustering with Multi-Modal LLMs 28 Mar 2025 · 0 repositories · arXiv:2503.22241
-
Enhance Generation Quality of Flow Matching V2A Model via Multi-Step CoT-Like Guidance and Combined Preference Optimization 28 Mar 2025 · 1 repository · arXiv:2503.22200
-
FLIP: Towards Comprehensive and Reliable Evaluation of Federated Prompt Learning 28 Mar 2025 · 1 repository · arXiv:2503.22263
-
Instance-Level Data-Use Auditing of Visual ML Models 28 Mar 2025 · 0 repositories · arXiv:2503.22413
-
SCHNet: SAM Marries CLIP for Human Parsing 28 Mar 2025 · 0 repositories · arXiv:2503.22237
-
Segment then Splat: A Unified Approach for 3D Open-Vocabulary Segmentation based on Gaussian Splatting 28 Mar 2025 · 0 repositories · arXiv:2503.22204
-
VisTa: Visual-contextual and Text-augmented Zero-shot Object-level OOD Detection 28 Mar 2025 · 0 repositories · arXiv:2503.22291
-
Semantic Library Adaptation: LoRA Retrieval and Fusion for Open-Vocabulary Semantic Segmentation 27 Mar 2025 · 1 repository · arXiv:2503.21780
-
Hierarchical Label Propagation: A Model-Size-Dependent Performance Booster for AudioSet Tagging 26 Mar 2025 · 1 repository · arXiv:2503.21826
-
Rethinking Vision-Language Model in Face Forensics: Multi-Modal Interpretable Forged Face Detector 26 Mar 2025 · 1 repository · arXiv:2503.20188Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 1 violated, 7 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 11 harvested samples)
-
Video Motion Graphs 26 Mar 2025 · 0 repositories · arXiv:2503.20218
-
VideoGEM: Training-free Action Grounding in Videos 26 Mar 2025 · 0 repositories · arXiv:2503.20348
-
Exploring Semantic Feature Discrimination for Perceptual Image Super-Resolution and Opinion-Unaware No-Reference Image Quality Assessment 25 Mar 2025 · 1 repository · arXiv:2503.19295Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
fine-CLIP: Enhancing Zero-Shot Fine-Grained Surgical Action Recognition with Vision-Language Models 25 Mar 2025 · 0 repositories · arXiv:2503.19670
-
Reverse Prompt: Cracking the Recipe Inside Text-to-Image Generation 25 Mar 2025 · 0 repositories · arXiv:2503.19937
-
Scaling Down Text Encoders of Text-to-Image Diffusion Models 25 Mar 2025 · 1 repository · arXiv:2503.19897
-
Enhanced OoD Detection through Cross-Modal Alignment of Multi-Modal Representations 24 Mar 2025 · 1 repository · arXiv:2503.18817
-
OCRT: Boosting Foundation Models in the Open World with Object-Concept-Relation Triad 24 Mar 2025 · 1 repository · arXiv:2503.18695
-
Panorama Generation From NFoV Image Done Right 24 Mar 2025 · 1 repository · arXiv:2503.18420
-
CustomKD: Customizing Large Vision Foundation for Edge Model Improvement via Knowledge Distillation 23 Mar 2025 · 0 repositories · arXiv:2503.18244
-
Text-Driven Cross-Modal Place Recognition Method for Remote Sensing Localization 23 Mar 2025 · 0 repositories · arXiv:2503.18035
-
GOAL: Global-local Object Alignment Learning 22 Mar 2025 · 1 repository · arXiv:2503.17782Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Multi-modality Anomaly Segmentation on the Road 22 Mar 2025 · 1 repository · arXiv:2503.17712
-
TDRI: Two-Phase Dialogue Refinement and Co-Adaptation for Interactive Image Generation 22 Mar 2025 · 0 repositories · arXiv:2503.17669
-
Improving Acoustic Scene Classification with City Features 21 Mar 2025 · 0 repositories · arXiv:2503.16862
-
Debugging and Runtime Analysis of Neural Networks with VLMs (A Case Study) 21 Mar 2025 · 0 repositories · arXiv:2503.17416
-
Enhancing Product Search Interfaces with Sketch-Guided Diffusion and Language Agents 21 Mar 2025 · 0 repositories · arXiv:2504.08739
-
Meme Similarity and Emotion Detection using Multimodal Analysis 21 Mar 2025 · 0 repositories · arXiv:2503.17493
-
PE-CLIP: A Parameter-Efficient Fine-Tuning of Vision Language Models for Dynamic Facial Expression Recognition 21 Mar 2025 · 0 repositories · arXiv:2503.16945
-
Seeing What Matters: Empowering CLIP with Patch Generation-to-Selection 21 Mar 2025 · 0 repositories · arXiv:2503.17080
-
CausalCLIPSeg: Unlocking CLIP's Potential in Referring Medical Image Segmentation with Causal Intervention 20 Mar 2025 · 1 repository · arXiv:2503.15949
-
Cross-Modal and Uncertainty-Aware Agglomeration for Open-Vocabulary 3D Scene Understanding 20 Mar 2025 · 1 repository · arXiv:2503.16707Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
OSLoPrompt: Bridging Low-Supervision Challenges and Open-Set Domain Generalization in CLIP 20 Mar 2025 · 1 repository · arXiv:2503.16106
-
Probabilistic Prompt Distribution Learning for Animal Pose Estimation 20 Mar 2025 · 1 repository · arXiv:2503.16120
-
ScalingNoise: Scaling Inference-Time Search for Generating Infinite Videos 20 Mar 2025 · 0 repositories · arXiv:2503.16400
-
STOP: Integrated Spatial-Temporal Dynamic Prompting for Video Understanding 20 Mar 2025 · 1 repository · arXiv:2503.15973
-
UniCrossAdapter: Multimodal Adaptation of CLIP for Radiology Report Generation 20 Mar 2025 · 1 repository · arXiv:2503.15940
-
V-NAW: Video-based Noise-aware Adaptive Weighting for Facial Expression Recognition 20 Mar 2025 · 1 repository · arXiv:2503.15970
-
FP4DiT: Towards Effective Floating Point Quantization for Diffusion Transformers 19 Mar 2025 · 1 repository · arXiv:2503.15465Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 1 honoured, 0 violated, 1 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples)
-
Recover and Match: Open-Vocabulary Multi-Label Recognition through Knowledge-Constrained Optimal Transport 19 Mar 2025 · 1 repository · arXiv:2503.15337Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)