Methods › Computer Vision › Vision and Language Pre-Trained Models › CLIP › Papers, page 5
Contrastive Language-Image Pre-training
CLIP
Papers archive 2025-07-28
archive papers tagged: 3,094 · with a code link: 1,617 · where Syntology ran a sample: 649 (554 with a run with no instrument failure, 95 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (649 of 3,094 tagged: 554 with a run with no instrument failure, 95 where every run was a failure of Syntology's instrument)
Page 5 of 31: papers 401 to 500 of 3,094, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
AdaGC: Improving Training Stability for Large Language Model Pretraining 16 Feb 2025 · 0 repositories · arXiv:2502.11034
-
Faces of Fairness: Examining Bias in Facial Expression Recognition Datasets and Models 16 Feb 2025 · 0 repositories · arXiv:2502.11049
-
Multi-Faceted Multimodal Monosemanticity 16 Feb 2025 · 0 repositories · arXiv:2502.14888
-
Demographic User Modeling for Social Robotics with Multimodal Pre-trained Models 15 Feb 2025 · 0 repositories · arXiv:2502.10642
-
Occlusion-aware Text-Image-Point Cloud Pretraining for Open-World 3D Object Recognition 15 Feb 2025 · 0 repositories · arXiv:2502.10674
-
Classifier-free Guidance with Adaptive Scaling 14 Feb 2025 · 1 repository · arXiv:2502.10574
-
Designing a Conditional Prior Distribution for Flow-Based Generative Models 13 Feb 2025 · 0 repositories · arXiv:2502.09611
-
GAIA: A Global, Multi-modal, Multi-scale Vision-Language Dataset for Remote Sensing Image Analysis 13 Feb 2025 · 1 repository · arXiv:2502.09598
-
Long-Term TalkingFace Generation via Motion-Prior Conditional Diffusion Model 13 Feb 2025 · 0 repositories · arXiv:2502.09533
-
When and How Does CLIP Enable Domain and Compositional Generalization? 13 Feb 2025 · 0 repositories · arXiv:2502.09507Syntology 24 ran (of which 0 constructed an object rather than computing a result; 16 with no instrument failure: 0 honoured, 0 violated, 16 with no contract checked; 8 where Syntology's instrument failed) · 8 unverified (of 32 harvested samples) · 5 pointer-only (licence)
-
Skrr: Skip and Re-use Text Encoder Layers for Memory Efficient Text-to-Image Generation 12 Feb 2025 · 0 repositories · arXiv:2502.08690
-
Captured by Captions: On Memorization and its Mitigation in CLIP Models 11 Feb 2025 · 0 repositories · arXiv:2502.07830
-
Direct Ascent Synthesis: Revealing Hidden Generative Capabilities in Discriminative Models 11 Feb 2025 · 1 repository · arXiv:2502.07753
-
Intrinsic Bias is Predicted by Pretraining Data and Correlates with Downstream Performance in Vision-Language Encoders 11 Feb 2025 · 1 repository · arXiv:2502.07957
-
MGPATH: Vision-Language Model with Multi-Granular Prompt Learning for Few-Shot WSI Classification 11 Feb 2025 · 1 repository · arXiv:2502.07409
-
Playmate: Flexible Control of Portrait Animation via 3D-Implicit Space Guided Diffusion 11 Feb 2025 · 0 repositories · arXiv:2502.07203
-
Scaling Pre-training to One Hundred Billion Data for Vision Language Models 11 Feb 2025 · 0 repositories · arXiv:2502.07617
-
Group-CLIP Uncertainty Modeling for Group Re-Identification 10 Feb 2025 · 0 repositories · arXiv:2502.06460
-
UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding 8 Feb 2025 · 1 repository · arXiv:2502.05415
-
Bridging Scales in Map Generation: A scale-aware cascaded generative mapping framework for seamless and consistent multi-scale cartographic representation 7 Feb 2025 · 0 repositories · arXiv:2502.04991
-
Shifting Attention to You: Personalized Brain-Inspired AI Models 7 Feb 2025 · 0 repositories · arXiv:2502.04658
-
Color in Visual-Language Models: CLIP deficiencies 6 Feb 2025 · 0 repositories · arXiv:2502.04470
-
ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features 6 Feb 2025 · 1 repository · arXiv:2502.04320Syntology official (archive's flag): 4 ran · 4 ran (of which 2 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 5 unverified (of 9 harvested samples) · 9 pointer-only (licence)
-
Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion 6 Feb 2025 · 1 repository · arXiv:2502.04263
-
Efficient Few-Shot Continual Learning in Vision-Language Models 6 Feb 2025 · 0 repositories · arXiv:2502.04098
-
CLIP Behaves like a Bag-of-Words Model Cross-modally but not Uni-modally 5 Feb 2025 · 1 repository · arXiv:2502.03566Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Disentangling CLIP for Multi-Object Perception 5 Feb 2025 · 0 repositories · arXiv:2502.02977
-
Rethinking the Global Knowledge of CLIP in Training-Free Open-Vocabulary Semantic Segmentation 5 Feb 2025 · 0 repositories · arXiv:2502.06818
-
Kronecker Mask and Interpretive Prompts are Language-Action Video Learners 5 Feb 2025 · 1 repository · arXiv:2502.03549Syntology official (archive's flag): 9 ran · 9 ran (of which 9 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified; every one of the 9 samples that ran constructed an object rather than computing a result (of 11 harvested samples) · 11 pointer-only (licence)
-
TexLiDAR: Automated Text Understanding for Panoramic LiDAR Data 5 Feb 2025 · 1 repository · arXiv:2502.04385
-
LoRA-TTT: Low-Rank Test-Time Training for Vision-Language Models 4 Feb 2025 · 0 repositories · arXiv:2502.02069
-
Rethinking Homogeneity of Vision and Text Tokens in Large Vision-and-Language Models 4 Feb 2025 · 0 repositories · arXiv:2502.01906
-
CLIP-DQA: Blindly Evaluating Dehazed Images from Global and Local Perspectives Using CLIP 3 Feb 2025 · 1 repository · arXiv:2502.01707
-
CLIP-UP: A Simple and Efficient Mixture-of-Experts CLIP Training Recipe with Sparse Upcycling 3 Feb 2025 · 0 repositories · arXiv:2502.00965
-
Detecting Backdoor Samples in Contrastive Language Image Pretraining 3 Feb 2025 · 1 repository · arXiv:2502.01385Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Robust-LLaVA: On the Effectiveness of Large-Scale Robust Image Encoders for Multi-modal Large Language Models 3 Feb 2025 · 1 repository · arXiv:2502.01576
-
PhiP-G: Physics-Guided Text-to-3D Compositional Scene Generation 2 Feb 2025 · 0 repositories · arXiv:2502.00708
-
UniGraph2: Learning a Unified Embedding Space to Bind Multimodal Graphs 2 Feb 2025 · 1 repository · arXiv:2502.00806
-
Leveraging Stable Diffusion for Monocular Depth Estimation via Image Semantic Encoding 1 Feb 2025 · 0 repositories · arXiv:2502.01666
-
ALBAR: Adversarial Learning approach to mitigate Biases in Action Recognition 31 Jan 2025 · 0 repositories · arXiv:2502.00156
-
Contrast-Aware Calibration for Fine-Tuned CLIP: Leveraging Image-Text Alignment 31 Jan 2025 · 0 repositories · arXiv:2501.19060
-
Fairness Analysis of CLIP-Based Foundation Models for X-Ray Image Classification 31 Jan 2025 · 0 repositories · arXiv:2501.19086
-
Laser: Efficient Language-Guided Segmentation in Neural Radiance Fields 31 Jan 2025 · 1 repository · arXiv:2501.19084
-
Lifting by Gaussians: A Simple, Fast and Flexible Method for 3D Instance Segmentation 31 Jan 2025 · 0 repositories · arXiv:2502.00173
-
Advances in Multimodal Adaptation and Generalization: From Traditional Approaches to Foundation Models 30 Jan 2025 · 5 repositories · arXiv:2501.18592
-
Efficient Redundancy Reduction for Open-Vocabulary Semantic Segmentation 29 Jan 2025 · 1 repository · arXiv:2501.17642
-
Technical report on label-informed logit redistribution for better domain generalization in low-shot classification with foundation models 29 Jan 2025 · 0 repositories · arXiv:2501.17595
-
Modulating CNN Features with Pre-Trained ViT Representations for Open-Vocabulary Object Detection 28 Jan 2025 · 0 repositories · arXiv:2501.16981
-
One Head Eight Arms: Block Matrix based Low Rank Adaptation for CLIP-based Few-Shot Learning 28 Jan 2025 · 0 repositories · arXiv:2501.16720
-
CILP-FGDI: Exploiting Vision-Language Model for Generalizable Person Re-Identification 27 Jan 2025 · 1 repository · arXiv:2501.16065
-
SPECIAL: Zero-shot Hyperspectral Image Classification With CLIP 27 Jan 2025 · 1 repository · arXiv:2501.16222
-
Domain Adaptation from Generated Multi-Weather Images for Unsupervised Maritime Object Classification 26 Jan 2025 · 0 repositories · arXiv:2501.15503
-
Fine Tuning without Catastrophic Forgetting via Selective Low Rank Adaptation 26 Jan 2025 · 0 repositories · arXiv:2501.15377
-
A Training-free Synthetic Data Selection Method for Semantic Segmentation 25 Jan 2025 · 1 repository · arXiv:2501.15201
-
Enhancing Intent Understanding for Ambiguous prompt: A Human-Machine Co-Adaption Strategy 25 Jan 2025 · 0 repositories · arXiv:2501.15167
-
Large-scale and Fine-grained Vision-language Pre-training for Enhanced CT Image Understanding 24 Jan 2025 · 1 repository · arXiv:2501.14548
-
Attribute-based Visual Reprogramming for Image Classification with CLIP 23 Jan 2025 · 1 repository · arXiv:2501.13982
-
EventVL: Understand Event Streams via Multimodal Large Language Model 23 Jan 2025 · 0 repositories · arXiv:2501.13707
-
Language modulates vision: Evidence from neural networks and human brain-lesion models 23 Jan 2025 · 0 repositories · arXiv:2501.13628
-
Large Vision-Language Models for Knowledge-Grounded Data Annotation of Memes 23 Jan 2025 · 1 repository · arXiv:2501.13851
-
Meta-Feature Adapter: Integrating Environmental Metadata for Enhanced Animal Re-identification 23 Jan 2025 · 0 repositories · arXiv:2501.13368
-
Text-driven Online Action Detection 23 Jan 2025 · 1 repository · arXiv:2501.13518
-
Accelerate High-Quality Diffusion Models with Inner Loop Feedback 22 Jan 2025 · 0 repositories · arXiv:2501.13107
-
Adapting OpenAI's CLIP Model for Few-Shot Image Inspection in Manufacturing Quality Control: An Expository Case Study with Multiple Application Examples 22 Jan 2025 · 0 repositories · arXiv:2501.12596
-
Can masking background and object reduce static bias for zero-shot action recognition? 22 Jan 2025 · 0 repositories · arXiv:2501.12681
-
A Comprehensive Social Bias Audit of Contrastive Vision Language Models 22 Jan 2025 · 0 repositories · arXiv:2501.13223
-
TeD-Loc: Text Distillation for Weakly Supervised Object Localization 22 Jan 2025 · 1 repository · arXiv:2501.12632
-
Audio Texture Manipulation by Exemplar-Based Analogy 21 Jan 2025 · 0 repositories · arXiv:2501.12385
-
SplitQuant: Layer Splitting for Low-Bit Neural Network Quantization 21 Jan 2025 · 0 repositories · arXiv:2501.12428
-
CatV2TON: Taming Diffusion Transformers for Vision-Based Virtual Try-On with Temporal Concatenation 20 Jan 2025 · 1 repository · arXiv:2501.11325
-
KPL: Training-Free Medical Knowledge Mining of Vision-Language Models 20 Jan 2025 · 1 repository · arXiv:2501.11231
-
Know "No'' Better: A Data-Driven Approach for Enhancing Negation Awareness in CLIP 19 Jan 2025 · 0 repositories · arXiv:2501.10913
-
ProKeR: A Kernel Perspective on Few-Shot Adaptation of Large Vision-Language Models 19 Jan 2025 · 0 repositories · arXiv:2501.11175
-
Efficient Auto-Labeling of Large-Scale Poultry Datasets (ALPD) Using Semi-Supervised Models, Active Learning, and Prompt-then-Detect Approach 18 Jan 2025 · 0 repositories · arXiv:2501.10809
-
CLIP-PCQA: Exploring Subjective-Aligned Vision-Language Modeling for Point Cloud Quality Assessment 17 Jan 2025 · 1 repository · arXiv:2501.10071
-
AnyStory: Towards Unified Single and Multiple Subject Personalization in Text-to-Image Generation 16 Jan 2025 · 1 repository · arXiv:2501.09503
-
Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness 16 Jan 2025 · 0 repositories · arXiv:2501.09446
-
On Learning Informative Trajectory Embeddings for Imitation, Classification and Regression 16 Jan 2025 · 1 repository · arXiv:2501.09327
-
Vision-Language Models Do Not Understand Negation 16 Jan 2025 · 0 repositories · arXiv:2501.09425
-
Benchmarking Robustness of Contrastive Learning Models for Medical Image-Report Retrieval 15 Jan 2025 · 0 repositories · arXiv:2501.09134
-
CityLoc: 6DoF Pose Distributional Localization for Text Descriptions in Large-Scale Scenes with Gaussian Representation 15 Jan 2025 · 0 repositories · arXiv:2501.08982
-
IDEA: Image Description Enhanced CLIP-Adapter 15 Jan 2025 · 1 repository · arXiv:2501.08816
-
Learning to Adapt Frozen CLIP for Few-Shot Test-Time Domain Adaptation 15 Jan 2025 · 1 repository
-
SHYI: Action Support for Contrastive Learning in High-Fidelity Text-to-Image Generation 15 Jan 2025 · 0 repositories · arXiv:2501.09055
-
Cross-Modal Transferable Image-to-Video Attack on Video Quality Metrics 14 Jan 2025 · 1 repository · arXiv:2501.08415
-
FLAVARS: A Multimodal Foundational Language and Vision Alignment Model for Remote Sensing 14 Jan 2025 · 0 repositories · arXiv:2501.08490
-
Uncovering Bias in Foundation Models: Impact, Testing, Harm, and Mitigation 14 Jan 2025 · 1 repository · arXiv:2501.10453
-
Exploring the Use of Contrastive Language-Image Pre-Training for Human Posture Classification: Insights from Yoga Pose Analysis 13 Jan 2025 · 0 repositories · arXiv:2501.07221
-
SST-EM: Advanced Metrics for Evaluating Semantic, Spatial and Temporal Aspects in Video Editing 13 Jan 2025 · 1 repository · arXiv:2501.07554
-
Evaluating Sample Utility for Data Selection by Mimicking Model Weights 12 Jan 2025 · 0 repositories · arXiv:2501.06708
-
MedGrad E-CLIP: Enhancing Trust and Transparency in AI-Driven Skin Lesion Diagnosis 12 Jan 2025 · 0 repositories · arXiv:2501.06887
-
RSRefSeg: Referring Remote Sensing Image Segmentation with Foundation Models 12 Jan 2025 · 1 repository · arXiv:2501.06809
-
Semantic-CD: Remote Sensing Image Semantic Change Detection towards Open-vocabulary Setting 12 Jan 2025 · 0 repositories · arXiv:2501.06808
-
A Holistically Point-guided Text Framework for Weakly-Supervised Camouflaged Object Detection 10 Jan 2025 · 0 repositories · arXiv:2501.06038
-
Generate, Transduct, Adapt: Iterative Transduction with VLMs 10 Jan 2025 · 0 repositories · arXiv:2501.06031
-
Multi-subject Open-set Personalization in Video Generation 10 Jan 2025 · 0 repositories · arXiv:2501.06187
-
StarGen: A Spatiotemporal Autoregression Framework with Video Diffusion Model for Scalable and Controllable Scene Generation 10 Jan 2025 · 0 repositories · arXiv:2501.05763
-
Discovering Hidden Visual Concepts Beyond Linguistic Input in Infant Learning 9 Jan 2025 · 1 repository · arXiv:2501.05205
-
Harnessing Large Language and Vision-Language Models for Robust Out-of-Distribution Detection 9 Jan 2025 · 0 repositories · arXiv:2501.05228
-
Mechanistic understanding and validation of large AI models with SemanticLens 9 Jan 2025 · 1 repository · arXiv:2501.05398Syntology official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 10 harvested samples)