Methods › Computer Vision › Vision and Language Pre-Trained Models › CLIP › Papers, page 16
Contrastive Language-Image Pre-training
CLIP
Papers archive 2025-07-28
archive papers tagged: 3,094 · with a code link: 1,617 · where Syntology ran a sample: 649 (554 with a run with no instrument failure, 95 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (649 of 3,094 tagged: 554 with a run with no instrument failure, 95 where every run was a failure of Syntology's instrument)
Page 16 of 31: papers 1,501 to 1,600 of 3,094, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Online Embedding Multi-Scale CLIP Features into 3D Maps 27 Mar 2024 · 0 repositories · arXiv:2403.18178
-
TextCraftor: Your Text Encoder Can be Image Quality Controller 27 Mar 2024 · 0 repositories · arXiv:2403.18978
-
Dual Memory Networks: A Versatile Adaptation Approach for Vision-Language Models 26 Mar 2024 · 1 repository · arXiv:2403.17589Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 4 where Syntology's instrument failed) · 3 unverified (of 10 harvested samples) · 4 pointer-only (licence)
-
OmniVid: A Generative Framework for Universal Video Understanding 26 Mar 2024 · 1 repository · arXiv:2403.17935Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 1 where Syntology's instrument failed) · 4 unverified (of 11 harvested samples) · 11 pointer-only (licence)
-
WordRobe: Text-Guided Generation of Textured 3D Garments 26 Mar 2024 · 0 repositories · arXiv:2403.17541
-
An Intermediate Fusion ViT Enables Efficient Text-Image Alignment in Diffusion Models 25 Mar 2024 · 0 repositories · arXiv:2403.16530
-
Continuous, Subject-Specific Attribute Control in T2I Models by Identifying Semantic Directions 25 Mar 2024 · 1 repository · arXiv:2403.17064Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples)
-
DreamLIP: Language-Image Pre-training with Long Captions 25 Mar 2024 · 1 repository · arXiv:2403.17007Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 1 violated, 2 with no contract checked; 2 where Syntology's instrument failed) · 3 unverified (of 8 harvested samples) · 8 pointer-only (licence)
-
Make-Your-Anchor: A Diffusion-based 2D Avatar Generation Framework 25 Mar 2024 · 1 repository · arXiv:2403.16510Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 2 where Syntology's instrument failed) · 2 unverified (of 8 harvested samples) · 8 pointer-only (licence)
-
Task2Box: Box Embeddings for Modeling Asymmetric Task Relationships 25 Mar 2024 · 1 repository · arXiv:2403.17173Syntology official (archive's flag): 1 ran · 1 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified; the one sample that ran constructed an object rather than computing a result (of 1 harvested sample) · 1 pointer-only (licence)
-
Centered Masking for Language-Image Pre-Training 23 Mar 2024 · 1 repository · arXiv:2403.15837
-
CLIP-VQDiffusion : Langauge Free Training of Text To Image generation using CLIP and vector quantized diffusion model 22 Mar 2024 · 1 repository · arXiv:2403.14944
-
FairerCLIP: Debiasing CLIP's Zero-Shot Predictions using Functions in RKHSs 22 Mar 2024 · 0 repositories · arXiv:2403.15593
-
LeGO: Leveraging a Surface Deformation Network for Animatable Stylized Face Generation with One Example 22 Mar 2024 · 1 repository · arXiv:2403.15227
-
LLaVA-PruMerge: Adaptive Token Reduction for Efficient Large Multimodal Models 22 Mar 2024 · 1 repository · arXiv:2403.15388Syntology 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Long-CLIP: Unlocking the Long-Text Capability of CLIP 22 Mar 2024 · 1 repository · arXiv:2403.15378Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 1 violated, 4 with no contract checked; 3 where Syntology's instrument failed) · 7 unverified (of 15 harvested samples) · 10 pointer-only (licence)
-
Transfer CLIP for Generalizable Image Denoising 22 Mar 2024 · 1 repository · arXiv:2403.15132
-
C-TPT: Calibrated Test-Time Prompt Tuning for Vision-Language Models via Text Feature Dispersion 21 Mar 2024 · 1 repository · arXiv:2403.14119Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 6 where Syntology's instrument failed) · 2 unverified (of 10 harvested samples) · 5 pointer-only (licence)
-
CFPL-FAS: Class Free Prompt Learning for Generalizable Face Anti-spoofing 21 Mar 2024 · 0 repositories · arXiv:2403.14333
-
Lexicon-Level Contrastive Visual-Grounding Improves Language Modeling 21 Mar 2024 · 2 repositories · arXiv:2403.14551
-
Locating and Mitigating Gender Bias in Large Language Models 21 Mar 2024 · 0 repositories · arXiv:2403.14409
-
OTSeg: Multi-prompt Sinkhorn Attention for Zero-Shot Semantic Segmentation 21 Mar 2024 · 1 repository · arXiv:2403.14183Syntology official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 1 honoured, 0 violated, 6 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 10 harvested samples) · 10 pointer-only (licence)
-
Unified Static and Dynamic Network: Efficient Temporal Filtering for Video Grounding 21 Mar 2024 · 1 repository · arXiv:2403.14174
-
CLIPSwarm: Generating Drone Shows from Text Prompts with Vision-Language Models 20 Mar 2024 · 0 repositories · arXiv:2403.13467
-
Diffusion-based Human Motion Style Transfer with Semantic Guidance 20 Mar 2024 · 0 repositories · arXiv:2405.06646
-
FissionFusion: Fast Geometric Generation and Hierarchical Souping for Medical Image Analysis 20 Mar 2024 · 1 repository · arXiv:2403.13341
-
RAR: Retrieving And Ranking Augmented MLLMs for Visual Recognition 20 Mar 2024 · 2 repositories · arXiv:2403.13805Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 2 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 2 unverified (of 7 harvested samples) · 3 pointer-only (licence)
-
Adapting Visual-Language Models for Generalizable Anomaly Detection in Medical Images 19 Mar 2024 · 1 repository · arXiv:2403.12570Syntology official (archive's flag): 9 ran · 9 ran (of which 1 constructed an object rather than computing a result; 5 with no instrument failure: 1 honoured, 1 violated, 3 with no contract checked; 4 where Syntology's instrument failed) · 6 unverified (of 15 harvested samples) · 6 pointer-only (licence)
-
AnySkill: Learning Open-Vocabulary Physical Skill for Interactive Agents 19 Mar 2024 · 0 repositories · arXiv:2403.12835
-
As Firm As Their Foundations: Can open-sourced foundation models be used to create adversarial examples for downstream tasks? 19 Mar 2024 · 0 repositories · arXiv:2403.12693
-
Better Call SAL: Towards Learning to Segment Anything in Lidar 19 Mar 2024 · 1 repository · arXiv:2403.13129
-
CLIP-VIS: Adapting CLIP for Open-Vocabulary Video Instance Segmentation 19 Mar 2024 · 1 repository · arXiv:2403.12455
-
Arc2Face: A Foundation Model for ID-Consistent Human Faces 18 Mar 2024 · 3 repositories · arXiv:2403.11641Syntology community repositories only · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 4 harvested samples) · 3 pointer-only (licence)
-
Boosting Continual Learning of Vision-Language Models via Mixture-of-Experts Adapters 18 Mar 2024 · 2 repositories · arXiv:2403.11549Syntology official (archive's flag): 2 ran · 2 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Boosting Zero-Shot Human-Object Interaction Detection with Vision-Language Transfer 18 Mar 2024 · 1 repository
-
Data-Efficient Contrastive Language-Image Pretraining: Prioritizing Data Quality over Quantity 18 Mar 2024 · 1 repository · arXiv:2403.12267
-
A Sober Look at the Robustness of CLIPs to Spurious Features 18 Mar 2024 · 0 repositories · arXiv:2403.11497
-
End-to-end multi-modal product matching in fashion e-commerce 18 Mar 2024 · 0 repositories · arXiv:2403.11593
-
Meta-Prompting for Automating Zero-shot Visual Recognition with LLMs 18 Mar 2024 · 1 repository · arXiv:2403.11755Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 3 where Syntology's instrument failed) · 2 unverified (of 7 harvested samples) · 3 pointer-only (licence)
-
N-Modal Contrastive Losses with Applications to Social Media Data in Trimodal Space 18 Mar 2024 · 0 repositories · arXiv:2403.12747
-
MindEye2: Shared-Subject Models Enable fMRI-To-Image With 1 Hour of Data 17 Mar 2024 · 1 repository · arXiv:2403.11207Syntology official (archive's flag): 15 ran · 15 ran (of which 0 constructed an object rather than computing a result; 14 with no instrument failure: 0 honoured, 3 violated, 11 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 15 harvested samples) · 2 pointer-only (licence)
-
Quality-Aware Image-Text Alignment for Real-World Image Quality Assessment 17 Mar 2024 · 1 repository · arXiv:2403.11176Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 2 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
TAG: Guidance-free Open-Vocabulary Semantic Segmentation 17 Mar 2024 · 1 repository · arXiv:2403.11197
-
LuoJiaHOG: A Hierarchy Oriented Geo-aware Image Caption Dataset for Remote Sensing Image-Text Retrival 16 Mar 2024 · 0 repositories · arXiv:2403.10887
-
N2F2: Hierarchical Scene Understanding with Nested Neural Feature Fields 16 Mar 2024 · 0 repositories · arXiv:2403.10997
-
Benchmarking Zero-Shot Robustness of Multimodal Foundation Models: A Pilot Study 15 Mar 2024 · 1 repository · arXiv:2403.10499
-
CoLeCLIP: Open-Domain Continual Learning via Joint Task Prompt and Vocabulary Learning 15 Mar 2024 · 1 repository · arXiv:2403.10245Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples) · 1 pointer-only (licence)
-
Does the Performance of Text-to-Image Retrieval Models Generalize Beyond Captions-as-a-Query? 15 Mar 2024 · 1 repository
-
E4C: Enhance Editability for Text-Based Image Editing by Harnessing Efficient CLIP Guidance 15 Mar 2024 · 0 repositories · arXiv:2403.10133
-
Enhancing Human-Centered Dynamic Scene Understanding via Multiple LLMs Collaborated Reasoning 15 Mar 2024 · 0 repositories · arXiv:2403.10107
-
Unlocking the Multi-modal Potential of CLIP for Generalized Category Discovery 15 Mar 2024 · 1 repository · arXiv:2403.09974Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 5 where Syntology's instrument failed) · 7 unverified (of 14 harvested samples) · 3 pointer-only (licence)
-
Improving Medical Multi-modal Contrastive Learning with Expert Annotations 15 Mar 2024 · 1 repository · arXiv:2403.10153Syntology official (archive's flag): 5 ran · 5 ran (of which 2 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 2 where Syntology's instrument failed) · 7 unverified (of 12 harvested samples) · 12 pointer-only (licence)
-
Isotropic3D: Image-to-3D Generation Based on a Single CLIP Embedding 15 Mar 2024 · 1 repository · arXiv:2403.10395Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 3 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Leveraging vision-language models for fair facial attribute classification 15 Mar 2024 · 0 repositories · arXiv:2403.10624
-
Multiscale Matching Driven by Cross-Modal Similarity Consistency for Audio-Text Retrieval 15 Mar 2024 · 0 repositories · arXiv:2403.10146
-
Learning Spatiotemporal Inconsistency via Thumbnail Layout for Face Deepfake Detection 15 Mar 2024 · 2 repositories · arXiv:2403.10261
-
Annotation Free Semantic Segmentation with Vision Foundation Models 14 Mar 2024 · 0 repositories · arXiv:2403.09307
-
Anomaly Detection by Adapting a pre-trained Vision Language Model 14 Mar 2024 · 0 repositories · arXiv:2403.09493
-
CLIP-EBC: CLIP Can Count Accurately through Enhanced Blockwise Classification 14 Mar 2024 · 1 repository · arXiv:2403.09281
-
PosSAM: Panoptic Open-vocabulary Segment Anything 14 Mar 2024 · 1 repository · arXiv:2403.09620
-
Robust Light-Weight Facial Affective Behavior Recognition with CLIP 14 Mar 2024 · 1 repository · arXiv:2403.09915
-
The First to Know: How Token Distributions Reveal Hidden Knowledge in Large Vision-Language Models? 14 Mar 2024 · 1 repository · arXiv:2403.09037
-
XCoOp: Explainable Prompt Learning for Computer-Aided Diagnosis via Concept-guided Context Optimization 14 Mar 2024 · 0 repositories · arXiv:2403.09410
-
A Multimodal Fusion Network For Student Emotion Recognition Based on Transformer and Tensor Product 13 Mar 2024 · 0 repositories · arXiv:2403.08511
-
Language-Driven Visual Consensus for Zero-Shot Semantic Segmentation 13 Mar 2024 · 0 repositories · arXiv:2403.08426
-
Robust COVID-19 Detection in CT Images with CLIP 13 Mar 2024 · 1 repository · arXiv:2403.08947
-
Beyond Text: Frozen Large Language Models in Visual Signal Comprehension 12 Mar 2024 · 1 repository · arXiv:2403.07874Syntology official (archive's flag): 17 ran · 17 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 1 honoured, 0 violated, 5 with no contract checked; 11 where Syntology's instrument failed) · 9 unverified (of 26 harvested samples) · 26 pointer-only (licence)
-
Calibrating Multi-modal Representations: A Pursuit of Group Robustness without Annotations 12 Mar 2024 · 1 repository · arXiv:2403.07241Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 3 harvested samples)
-
Learning Generalizable Feature Fields for Mobile Manipulation 12 Mar 2024 · 0 repositories · arXiv:2403.07563
-
MoPE-CLIP: Structured Pruning for Efficient Vision-Language Models with Module-wise Pruning Error Metric 12 Mar 2024 · 0 repositories · arXiv:2403.07839
-
Towards Zero-shot Human-Object Interaction Detection via Vision-Language Integration 12 Mar 2024 · 0 repositories · arXiv:2403.07246
-
Unified Source-Free Domain Adaptation 12 Mar 2024 · 1 repository · arXiv:2403.07601
-
You'll Never Walk Alone: A Sketch and Text Duet for Fine-Grained Image Retrieval 12 Mar 2024 · 0 repositories · arXiv:2403.07222
-
Boosting Image Restoration via Priors from Pre-trained Models 11 Mar 2024 · 0 repositories · arXiv:2403.06793
-
Human Pose Descriptions and Subject-Focused Attention for Improved Zero-Shot Transfer in Human-Centric Classification Tasks 11 Mar 2024 · 0 repositories · arXiv:2403.06904
-
FontCLIP: A Semantic Typography Visual-Language Model for Multilingual Font Applications 11 Mar 2024 · 1 repository · arXiv:2403.06453
-
Semantic Residual Prompts for Continual Learning 11 Mar 2024 · 1 repository · arXiv:2403.06870
-
Split to Merge: Unifying Separated Modalities for Unsupervised Domain Adaptation 11 Mar 2024 · 1 repository · arXiv:2403.06946Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 1 violated, 5 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 9 harvested samples) · 9 pointer-only (licence)
-
Style2Talker: High-Resolution Talking Head Generation with Emotion Style and Art Style 11 Mar 2024 · 0 repositories · arXiv:2403.06365
-
Toward Generalist Anomaly Detection via In-context Residual Learning with Few-shot Sample Prompts 11 Mar 2024 · 2 repositories · arXiv:2403.06495Syntology official (archive's flag): 16 ran · 16 ran (of which 0 constructed an object rather than computing a result; 14 with no instrument failure: 0 honoured, 1 violated, 13 with no contract checked; 2 where Syntology's instrument failed) · 5 unverified (of 21 harvested samples) · 16 pointer-only (licence)
-
In-context Prompt Learning for Test-time Vision Recognition with Frozen Vision-language Model 10 Mar 2024 · 0 repositories · arXiv:2403.06126
-
RESTORE: Towards Feature Shift for Vision-Language Prompt Learning 10 Mar 2024 · 1 repository · arXiv:2403.06136
-
Test-time Distribution Learning Adapter for Cross-modal Visual Reasoning 10 Mar 2024 · 0 repositories · arXiv:2403.06059
-
ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment 8 Mar 2024 · 2 repositories · arXiv:2403.05135Syntology 5 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 2 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples)
-
Exploring Robust Features for Few-Shot Object Detection in Satellite Imagery 8 Mar 2024 · 1 repository · arXiv:2403.05381
-
PEEB: Part-based Image Classifiers with an Explainable and Editable Language Bottleneck 8 Mar 2024 · 1 repository · arXiv:2403.05297Syntology official (archive's flag): 14 ran · 14 ran (of which 0 constructed an object rather than computing a result; 13 with no instrument failure: 0 honoured, 0 violated, 13 with no contract checked; 1 where Syntology's instrument failed) · 7 unverified (of 21 harvested samples)
-
A³lign-DFER: Pioneering Comprehensive Dynamic Affective Alignment for Dynamic Facial Expression Recognition with CLIP 7 Mar 2024 · 0 repositories · arXiv:2403.04294
-
CLIP the Bias: How Useful is Balancing Data in Multimodal Learning? 7 Mar 2024 · 0 repositories · arXiv:2403.04547
-
Self-Adapting Large Visual-Language Models to Edge Devices across Visual Modalities 7 Mar 2024 · 1 repository · arXiv:2403.04908Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Contrastive Learning of Person-independent Representations for Facial Action Unit Detection 6 Mar 2024 · 0 repositories · arXiv:2403.03400
-
FLAME Diffuser: Wildfire Image Synthesis using Mask Guided Diffusion 6 Mar 2024 · 1 repository · arXiv:2403.03463Syntology official (archive's flag): 13 ran · 13 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 6 where Syntology's instrument failed) · 3 unverified (of 16 harvested samples) · 16 pointer-only (licence)
-
MeaCap: Memory-Augmented Zero-shot Image Captioning 6 Mar 2024 · 1 repository · arXiv:2403.03715Syntology official (archive's flag): 6 ran · 6 ran (of which 1 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 2 where Syntology's instrument failed) · 4 unverified (of 10 harvested samples) · 10 pointer-only (licence)
-
Scene Depth Estimation from Traditional Oriental Landscape Paintings 6 Mar 2024 · 0 repositories · arXiv:2403.03408
-
CLEVR-POC: Reasoning-Intensive Visual Question Answering in Partially Observable Environments 5 Mar 2024 · 0 repositories · arXiv:2403.03203
-
DomainVerse: A Benchmark Towards Real-World Distribution Shifts For Tuning-Free Adaptive Domain Generalization 5 Mar 2024 · 0 repositories · arXiv:2403.02714
-
Finetuned Multimodal Language Models Are High-Quality Image-Text Data Filters 5 Mar 2024 · 0 repositories · arXiv:2403.02677
-
Modeling Collaborator: Enabling Subjective Vision Classification With Minimal Human Effort via LLM Tool-Use 5 Mar 2024 · 0 repositories · arXiv:2403.02626
-
PromptKD: Unsupervised Prompt Distillation for Vision-Language Models 5 Mar 2024 · 1 repository · arXiv:2403.02781
-
What do we learn from inverting CLIP models? 5 Mar 2024 · 1 repository · arXiv:2403.02580Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
FreeA: Human-object Interaction Detection using Free Annotation Labels 4 Mar 2024 · 0 repositories · arXiv:2403.01840