{"url":"/method/clip","slug":"clip","name":"CLIP","full_name":"Contrastive Language-Image Pre-training","full_name_withheld":false,"description_markdown":"**Contrastive Language-Image Pre-training** (**CLIP**), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning from natural language supervision. , CLIP jointly trains an image encoder and a text encoder to predict the correct pairings of a batch of (image, text) training examples. At test time the learned text encoder synthesizes a zero-shot linear classifier by embedding the names or descriptions of the target dataset’s classes. \r\n\r\nFor pre-training, CLIP is trained to predict which of the $N X N$ possible (image, text) pairings across a batch actually occurred. CLIP learns a multi-modal embedding space by jointly training an image encoder and text encoder to maximize the cosine similarity of the image and text embeddings of the $N$ real pairs in the batch while minimizing the cosine similarity of the embeddings of the $N^2 - N$ incorrect pairings. A symmetric cross entropy loss is optimized over these similarity scores. \r\n\r\nImage credit: [Learning Transferable Visual Models From Natural Language Supervision](https://arxiv.org/pdf/2103.00020.pdf)","description_state":"present","introduced_year":null,"introduced_by":{"title":"Learning Transferable Visual Models From Natural Language Supervision","paper":"/paper/learning-transferable-visual-models-from","first_author":"Alec Radford","n_authors":12,"url_abs":null,"archive_paper_url":"https://paperswithcode.com/paper/learning-transferable-visual-models-from"},"source":{"url":"https://arxiv.org/abs/2103.00020v1","title":"Learning Transferable Visual Models From Natural Language Supervision","url_on_a_paper_host":true},"code_snippet_url":"https://github.com/OpenAI/CLIP","code_snippet_url_on_a_code_host":true,"categories":[{"area":"Computer Vision","area_id":"computer-vision","collection":"Vision and Language Pre-Trained Models","url":"/methods/category/vision-and-language-pre-trained-models","pwc_aliases":[]},{"area":"Computer Vision","area_id":"computer-vision","collection":"Image Representations","url":"/methods/category/image-representations","pwc_aliases":[]}],"n_papers_tagged":3094,"archive_num_papers":3094,"papers_newest_first":[{"paper":null,"title":"Bridge Feature Matching and Cross-Modal Alignment with Mutual-filtering for Zero-shot Anomaly Detection","date":"2025-07-15","arxiv_id":"2507.11003","n_code_links":0,"syntology":null},{"paper":null,"title":"CATVis: Context-Aware Thought Visualization","date":"2025-07-15","arxiv_id":"2507.11522","n_code_links":0,"syntology":null},{"paper":"/paper/dearli-decoupled-enhancement-of-recognition","title":"DEARLi: Decoupled Enhancement of Recognition and Localization for Semi-supervised Panoptic Segmentation","date":"2025-07-14","arxiv_id":"2507.10118","n_code_links":1,"syntology":null},{"paper":"/paper/test-time-canonicalization-by-foundation","title":"Test-Time Canonicalization by Foundation Models for Robust Perception","date":"2025-07-14","arxiv_id":"2507.10375","n_code_links":1,"syntology":{"ran":5,"of":5,"unverified":0,"pointer_only":0}},{"paper":"/paper/text-visual-semantic-constrained-ai-generated","title":"Text-Visual Semantic Constrained AI-Generated Image Quality Assessment","date":"2025-07-14","arxiv_id":"2507.10432","n_code_links":1,"syntology":null},{"paper":null,"title":"Calibrated and Robust Foundation Models for Vision-Language and Medical Image Tasks Under Distribution Shift","date":"2025-07-12","arxiv_id":"2507.09222","n_code_links":0,"syntology":null},{"paper":null,"title":"A Neural Representation Framework with LLM-Driven Spatial Reasoning for Open-Vocabulary 3D Visual Grounding","date":"2025-07-09","arxiv_id":"2507.06719","n_code_links":0,"syntology":null},{"paper":"/paper/cultureclip-empowering-clip-with-cultural","title":"CultureCLIP: Empowering CLIP with Cultural Awareness through Synthetic Images and Contextualized Captions","date":"2025-07-08","arxiv_id":"2507.06210","n_code_links":1,"syntology":{"ran":0,"of":11,"unverified":11,"pointer_only":11}},{"paper":null,"title":"Integrated Structural Prompt Learning for Vision-Language Models","date":"2025-07-08","arxiv_id":"2507.05677","n_code_links":0,"syntology":null},{"paper":"/paper/rsrefseg-2-decoupling-referring-remote","title":"RSRefSeg 2: Decoupling Referring Remote Sensing Image Segmentation with Foundation Models","date":"2025-07-08","arxiv_id":"2507.06231","n_code_links":1,"syntology":null},{"paper":"/paper/semi-supervised-defect-detection-via","title":"Semi-Supervised Defect Detection via Conditional Diffusion and CLIP-Guided Noise Filtering","date":"2025-07-08","arxiv_id":"2507.05588","n_code_links":1,"syntology":null},{"paper":null,"title":"An analysis of vision-language models for fabric retrieval","date":"2025-07-07","arxiv_id":"2507.04735","n_code_links":0,"syntology":null},{"paper":"/paper/clip-guided-backdoor-defense-through-entropy","title":"CLIP-Guided Backdoor Defense through Entropy-Based Poisoned Dataset Separation","date":"2025-07-07","arxiv_id":"2507.05113","n_code_links":1,"syntology":null},{"paper":"/paper/pfedmma-personalized-federated-fine-tuning","title":"pFedMMA: Personalized Federated Fine-Tuning with Multi-Modal Adapter for Vision-Language Models","date":"2025-07-07","arxiv_id":"2507.05394","n_code_links":1,"syntology":{"ran":1,"of":11,"unverified":10,"pointer_only":11}},{"paper":"/paper/beyond-accuracy-metrics-that-uncover-what","title":"Beyond Accuracy: Metrics that Uncover What Makes a 'Good' Visual Descriptor","date":"2025-07-04","arxiv_id":"2507.03542","n_code_links":1,"syntology":null},{"paper":null,"title":"Helping CLIP See Both the Forest and the Trees: A Decomposition and Description Approach","date":"2025-07-04","arxiv_id":"2507.03458","n_code_links":0,"syntology":null},{"paper":null,"title":"Prompt Disentanglement via Language Guidance and Representation Alignment for Domain Generalization","date":"2025-07-03","arxiv_id":"2507.02288","n_code_links":0,"syntology":null},{"paper":null,"title":"VoteSplat: Hough Voting Gaussian Splatting for 3D Scene Understanding","date":"2025-06-28","arxiv_id":"2506.22799","n_code_links":0,"syntology":null},{"paper":null,"title":"Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs","date":"2025-06-27","arxiv_id":"2506.22139","n_code_links":0,"syntology":null},{"paper":null,"title":"Little By Little: Continual Learning via Self-Activated Sparse Mixture-of-Rank Adaptive Learning","date":"2025-06-26","arxiv_id":"2506.21035","n_code_links":0,"syntology":null},{"paper":"/paper/mitigating-hallucination-of-large-vision","title":"Mitigating Hallucination of Large Vision-Language Models via Dynamic Logits Calibration","date":"2025-06-26","arxiv_id":"2506.21509","n_code_links":1,"syntology":null},{"paper":null,"title":"Multimodal Prompt Alignment for Facial Expression Recognition","date":"2025-06-26","arxiv_id":"2506.21017","n_code_links":0,"syntology":null},{"paper":"/paper/reme-a-data-centric-framework-for-training","title":"ReME: A Data-Centric Framework for Training-Free Open-Vocabulary Segmentation","date":"2025-06-26","arxiv_id":"2506.21233","n_code_links":1,"syntology":null},{"paper":"/paper/sharpzo-hybrid-sharpness-aware-vision","title":"SharpZO: Hybrid Sharpness-Aware Vision Language Model Prompt Tuning via Forward-Only Passes","date":"2025-06-26","arxiv_id":"2506.20990","n_code_links":1,"syntology":{"ran":4,"of":10,"unverified":6,"pointer_only":0}},{"paper":null,"title":"Leveraging Vision-Language Models to Select Trustworthy Super-Resolution Samples Generated by Diffusion Models","date":"2025-06-25","arxiv_id":"2506.20832","n_code_links":0,"syntology":null},{"paper":null,"title":"Unfolding the Past: A Comprehensive Deep Learning Approach to Analyzing Incunabula Pages","date":"2025-06-22","arxiv_id":"2506.18069","n_code_links":0,"syntology":null},{"paper":null,"title":"Multimodal Political Bias Identification and Neutralization","date":"2025-06-20","arxiv_id":"2506.17372","n_code_links":0,"syntology":null},{"paper":null,"title":"Prmpt2Adpt: Prompt-Based Zero-Shot Domain Adaptation for Resource-Constrained Environments","date":"2025-06-20","arxiv_id":"2506.16994","n_code_links":0,"syntology":null},{"paper":null,"title":"Can Common VLMs Rival Medical VLMs? Evaluation and Strategic Insights","date":"2025-06-19","arxiv_id":"2506.17337","n_code_links":0,"syntology":null},{"paper":"/paper/evolutionary-caching-to-accelerate-your-off","title":"Evolutionary Caching to Accelerate Your Off-the-Shelf Diffusion Model","date":"2025-06-18","arxiv_id":"2506.15682","n_code_links":1,"syntology":{"ran":0,"of":1,"unverified":1,"pointer_only":1}}],"papers_shown":30,"tasks":[{"task":"/task/retrieval","name":"Retrieval","papers":336},{"task":"/task/language-modelling","name":"Language Modelling","papers":297},{"task":"/task/zero-shot-learning","name":"Zero-Shot Learning","papers":253},{"task":"/task/image-classification","name":"Image Classification","papers":250},{"task":"/task/image-generation","name":"Image Generation","papers":245},{"task":"/task/contrastive-learning","name":"Contrastive Learning","papers":228},{"task":"/task/semantic-segmentation","name":"Semantic Segmentation","papers":219},{"task":"/task/image-classification","name":"image-classification","papers":218},{"task":"/task/language-modeling","name":"Language Modeling","papers":193},{"task":"/task/prompt-learning","name":"Prompt Learning","papers":162},{"task":"/task/segmentation","name":"Segmentation","papers":161},{"task":null,"name":"zero-shot-classification","papers":158},{"task":"/task/representation-learning","name":"Representation Learning","papers":143},{"task":"/task/object","name":"Object","papers":142},{"task":"/task/image-captioning","name":"Image Captioning","papers":117},{"task":"/task/object-detection","name":"Object Detection","papers":117},{"task":"/task/object-detection-1","name":"object-detection","papers":115},{"task":"/task/text-to-image-generation","name":"Text-to-Image Generation","papers":109},{"task":"/task/transfer-learning","name":"Transfer Learning","papers":107},{"task":"/task/image-retrieval","name":"Image Retrieval","papers":106}],"tasks_shown":20,"n_tasks":947,"usage_by_year":[{"year":"2021","papers":101},{"year":"2022","papers":359},{"year":"2023","papers":887},{"year":"2024","papers":1191},{"year":"2025","papers":556}],"row_source":"methods_table","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/clip"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}