{"url":"/method/coop","slug":"coop","name":"CoOp","full_name":"Context Optimization","full_name_withheld":false,"description_markdown":"**CoOp**, or **Context Optimization**, is an automated prompt engineering method that avoids manual prompt tuning by modeling context words with continuous vectors that are end-to-end learned from data. The context could be shared among all classes or designed to be class-specific. During training, we simply minimize the prediction error using the cross-entropy loss with respect to the learnable context vectors, while keeping the pre-trained parameters fixed. The gradients can be back-propagated all the way through the text encoder, distilling the rich knowledge encoded in the parameters for learning task-relevant context.","description_state":"present","introduced_year":null,"introduced_by":{"title":"Learning to Prompt for Vision-Language Models","paper":"/paper/learning-to-prompt-for-vision-language-models","first_author":"Kaiyang Zhou","n_authors":4,"url_abs":null,"archive_paper_url":"https://paperswithcode.com/paper/learning-to-prompt-for-vision-language-models"},"source":{"url":"https://arxiv.org/abs/2109.01134v6","title":"Learning to Prompt for Vision-Language Models","url_on_a_paper_host":true},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"General","area_id":"general","collection":"Prompt Engineering","url":"/methods/category/prompt-engineering","pwc_aliases":[]}],"n_papers_tagged":30,"archive_num_papers":30,"papers_newest_first":[{"paper":"/paper/fa-forced-prompt-learning-of-vision-language","title":"FA: Forced Prompt Learning of Vision-Language Models for Out-of-Distribution Detection","date":"2025-07-06","arxiv_id":"2507.04511","n_code_links":1,"syntology":{"ran":3,"of":9,"unverified":6,"pointer_only":9}},{"paper":"/paper/mmrl-parameter-efficient-and-interaction","title":"MMRL++: Parameter-Efficient and Interaction-Aware Representation Learning for Vision-Language Models","date":"2025-05-15","arxiv_id":"2505.10088","n_code_links":1,"syntology":null},{"paper":"/paper/mmrl-multi-modal-representation-learning-for","title":"MMRL: Multi-Modal Representation Learning for Vision-Language Models","date":"2025-03-11","arxiv_id":"2503.08497","n_code_links":1,"syntology":{"ran":0,"of":1,"unverified":1,"pointer_only":0}},{"paper":null,"title":"Multi-Point Positional Insertion Tuning for Small Object Detection","date":"2024-12-24","arxiv_id":"2412.18090","n_code_links":0,"syntology":null},{"paper":null,"title":"PLPP: Prompt Learning with Perplexity Is Self-Distillation for Vision-Language Models","date":"2024-12-18","arxiv_id":"2412.15277","n_code_links":0,"syntology":null},{"paper":"/paper/textrefiner-internal-visual-feature-as","title":"TextRefiner: Internal Visual Feature as Efficient Refiner for Vision-Language Models Prompt Tuning","date":"2024-12-11","arxiv_id":"2412.08176","n_code_links":1,"syntology":null},{"paper":"/paper/transagent-transfer-vision-language","title":"TransAgent: Transfer Vision-Language Foundation Models with Heterogeneous Agent Collaboration","date":"2024-10-16","arxiv_id":"2410.12183","n_code_links":1,"syntology":{"ran":6,"of":8,"unverified":2,"pointer_only":0}},{"paper":null,"title":"FLIER: Few-shot Language Image Models Embedded with Latent Representations","date":"2024-10-10","arxiv_id":"2410.07648","n_code_links":0,"syntology":null},{"paper":"/paper/understanding-and-mitigating-miscalibration","title":"Understanding and Mitigating Miscalibration in Prompt Tuning for Vision-Language Models","date":"2024-10-03","arxiv_id":"2410.02681","n_code_links":1,"syntology":{"ran":4,"of":9,"unverified":5,"pointer_only":0}},{"paper":null,"title":"Enhancing Robustness of Vision-Language Models through Orthogonality Learning and Self-Regularization","date":"2024-07-11","arxiv_id":"2407.08374","n_code_links":0,"syntology":null},{"paper":null,"title":"Learning to Adapt Category Consistent Meta-Feature of CLIP for Few-Shot Classification","date":"2024-07-08","arxiv_id":"2407.05647","n_code_links":0,"syntology":null},{"paper":null,"title":"IntCoOp: Interpretability-Aware Vision-Language Prompt Tuning","date":"2024-06-19","arxiv_id":"2406.13683","n_code_links":0,"syntology":null},{"paper":"/paper/aapl-adding-attributes-to-prompt-learning-for","title":"AAPL: Adding Attributes to Prompt Learning for Vision-Language Models","date":"2024-04-25","arxiv_id":"2404.16804","n_code_links":1,"syntology":null},{"paper":"/paper/weak-distribution-detectors-lead-to-stronger","title":"Weak Distribution Detectors Lead to Stronger Generalizability of Vision-Language Prompt Tuning","date":"2024-03-31","arxiv_id":"2404.00603","n_code_links":1,"syntology":{"ran":4,"of":6,"unverified":2,"pointer_only":0}},{"paper":null,"title":"Concept-Guided Prompt Learning for Generalization in Vision-Language Models","date":"2024-01-15","arxiv_id":"2401.07457","n_code_links":0,"syntology":null},{"paper":null,"title":"Text-driven Prompt Generation for Vision-Language Models in Federated Learning","date":"2023-10-09","arxiv_id":"2310.06123","n_code_links":0,"syntology":null},{"paper":null,"title":"SwapPrompt: Test-Time Prompt Adaptation for Vision-Language Models","date":"2023-09-21","arxiv_id":null,"n_code_links":0,"syntology":null},{"paper":"/paper/pre-vision-language-prompt-learning-with","title":"PRE: Vision-Language Prompt Learning with Reparameterization Encoder","date":"2023-09-14","arxiv_id":"2309.07760","n_code_links":2,"syntology":null},{"paper":"/paper/language-models-as-black-box-optimizers-for","title":"Language Models as Black-Box Optimizers for Vision-Language Models","date":"2023-09-12","arxiv_id":"2309.05950","n_code_links":1,"syntology":{"ran":7,"of":11,"unverified":4,"pointer_only":11}},{"paper":null,"title":"Context-Aware Prompt Tuning for Vision-Language Model with Dual-Alignment","date":"2023-09-08","arxiv_id":"2309.04158","n_code_links":0,"syntology":null},{"paper":null,"title":"Unsupervised Prototype Adapter for Vision-Language Models","date":"2023-08-22","arxiv_id":"2308.11507","n_code_links":0,"syntology":null},{"paper":"/paper/mudpt-multi-modal-deep-symphysis-prompt","title":"MuDPT: Multi-modal Deep-symphysis Prompt Tuning for Large Pre-trained Vision-Language Models","date":"2023-06-20","arxiv_id":"2306.11400","n_code_links":1,"syntology":null},{"paper":"/paper/locoop-few-shot-out-of-distribution-detection-1","title":"LoCoOp: Few-Shot Out-of-Distribution Detection via Prompt Learning","date":"2023-06-02","arxiv_id":"2306.01293","n_code_links":2,"syntology":{"ran":5,"of":12,"unverified":7,"pointer_only":0}},{"paper":null,"title":"Task-Oriented Multi-Modal Mutual Leaning for Vision-Language Models","date":"2023-03-30","arxiv_id":"2303.17169","n_code_links":0,"syntology":null},{"paper":"/paper/understanding-and-mitigating-overfitting-in","title":"Understanding and Mitigating Overfitting in Prompt Tuning for Vision-Language Models","date":"2022-11-04","arxiv_id":"2211.02219","n_code_links":1,"syntology":null},{"paper":"/paper/prompt-tuning-with-soft-context-sharing-for","title":"Prompt Tuning with Soft Context Sharing for Vision-Language Models","date":"2022-08-29","arxiv_id":"2208.13474","n_code_links":1,"syntology":null},{"paper":"/paper/learning-to-compose-soft-prompts-for","title":"Learning to Compose Soft Prompts for Compositional Zero-Shot Learning","date":"2022-04-07","arxiv_id":"2204.03574","n_code_links":1,"syntology":{"ran":2,"of":2,"unverified":0,"pointer_only":0}},{"paper":"/paper/unsupervised-prompt-learning-for-vision","title":"Unsupervised Prompt Learning for Vision-Language Models","date":"2022-04-07","arxiv_id":"2204.03649","n_code_links":1,"syntology":null},{"paper":"/paper/conditional-prompt-learning-for-vision","title":"Conditional Prompt Learning for Vision-Language Models","date":"2022-03-10","arxiv_id":"2203.05557","n_code_links":12,"syntology":{"ran":4,"of":6,"unverified":2,"pointer_only":0}},{"paper":"/paper/learning-to-prompt-for-vision-language-models","title":"Learning to Prompt for Vision-Language Models","date":"2021-09-02","arxiv_id":"2109.01134","n_code_links":18,"syntology":{"ran":3,"of":12,"unverified":9,"pointer_only":0}}],"papers_shown":30,"tasks":[{"task":"/task/prompt-learning","name":"Prompt Learning","papers":14},{"task":"/task/prompt-engineering","name":"Prompt Engineering","papers":7},{"task":"/task/domain-generalization","name":"Domain Generalization","papers":6},{"task":"/task/transfer-learning","name":"Transfer Learning","papers":5},{"task":"/task/few-shot-learning","name":"Few-Shot Learning","papers":4},{"task":"/task/image-classification","name":"Image Classification","papers":3},{"task":"/task/ood-detection","name":"Out of Distribution (OOD) Detection","papers":3},{"task":"/task/representation-learning","name":"Representation Learning","papers":3},{"task":"/task/zero-shot-learning","name":"Zero-Shot Learning","papers":3},{"task":"/task/image-classification","name":"image-classification","papers":3},{"task":"/task/attribute","name":"Attribute","papers":2},{"task":"/task/data-augmentation","name":"Data Augmentation","papers":2},{"task":"/task/few-shot-image-classification","name":"Few-Shot Image Classification","papers":2},{"task":"/task/image-generation","name":"Image Generation","papers":2},{"task":"/task/object","name":"Object","papers":2},{"task":"/task/object-detection","name":"Object Detection","papers":2},{"task":"/task/out-of-distribution-detection","name":"Out-of-Distribution Detection","papers":2},{"task":"/task/object-detection-1","name":"object-detection","papers":2},{"task":"/task/compositional-zero-shot-learning","name":"Compositional Zero-Shot Learning","papers":1},{"task":"/task/federated-learning","name":"Federated Learning","papers":1}],"tasks_shown":20,"n_tasks":39,"usage_by_year":[{"year":"2021","papers":1},{"year":"2022","papers":5},{"year":"2023","papers":9},{"year":"2024","papers":12},{"year":"2025","papers":3}],"row_source":"methods_table","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/coop"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}