{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/learning-to-prompt-for-vision-language-models","title":"Learning to Prompt for Vision-Language Models","arxiv_id":"2109.01134","date":"2021-09-02","proceeding":null,"authors":["Kaiyang Zhou","Jingkang Yang","Chen Change Loy","Ziwei Liu"],"abstract":"Large pre-trained vision-language models like CLIP have shown great potential in learning representations that are transferable across a wide range of downstream tasks. Different from the traditional representation learning that is based mostly on discretized labels, vision-language pre-training aligns images and texts in a common feature space, which allows zero-shot transfer to a downstream task via prompting, i.e., classification weights are synthesized from natural language describing classes of interest. In this work, we show that a major challenge for deploying such models in practice is prompt engineering, which requires domain expertise and is extremely time-consuming -- one needs to spend a significant amount of time on words tuning since a slight change in wording could have a huge impact on performance. Inspired by recent advances in prompt learning research in natural language processing (NLP), we propose Context Optimization (CoOp), a simple approach specifically for adapting CLIP-like vision-language models for downstream image recognition. Concretely, CoOp models a prompt's context words with learnable vectors while the entire pre-trained parameters are kept fixed. To handle different image recognition tasks, we provide two implementations of CoOp: unified context and class-specific context. Through extensive experiments on 11 datasets, we demonstrate that CoOp requires as few as one or two shots to beat hand-crafted prompts with a decent margin and is able to gain significant improvements over prompt engineering with more shots, e.g., with 16 shots the average gain is around 15% (with the highest reaching over 45%). Despite being a learning-based approach, CoOp achieves superb domain generalization performance compared with the zero-shot model using hand-crafted prompts.","url_abs":"https://arxiv.org/abs/2109.01134v6","url_pdf":"https://arxiv.org/pdf/2109.01134v6.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"learning-to-prompt-for-vision-language-models","repo_url":"https://github.com/kaiyangzhou/coop","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"learning-to-prompt-for-vision-language-models","repo_url":"https://github.com/ArsenalCheng/Meta-Adapter","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"learning-to-prompt-for-vision-language-models","repo_url":"https://github.com/Gahyeonkim09/AAPL","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"learning-to-prompt-for-vision-language-models","repo_url":"https://github.com/ThomasWangY/2024-AAAI-HPT","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"learning-to-prompt-for-vision-language-models","repo_url":"https://github.com/Vill-Lab/2024-TIP-MetaPrompt","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"learning-to-prompt-for-vision-language-models","repo_url":"https://github.com/YangYongJin/APEX","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}},{"paper_slug":"learning-to-prompt-for-vision-language-models","repo_url":"https://github.com/azshue/TPT","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"learning-to-prompt-for-vision-language-models","repo_url":"https://github.com/farinamatteo/zero","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"learning-to-prompt-for-vision-language-models","repo_url":"https://github.com/healthx-lab/biomedcoop","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"learning-to-prompt-for-vision-language-models","repo_url":"https://github.com/hhenryd/tap","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"learning-to-prompt-for-vision-language-models","repo_url":"https://github.com/kaiyangzhou/on-device-dg","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"NOASSERTION"}},{"paper_slug":"learning-to-prompt-for-vision-language-models","repo_url":"https://github.com/kenomo/industrial-clip","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"learning-to-prompt-for-vision-language-models","repo_url":"https://github.com/mlvlab/dapt","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"learning-to-prompt-for-vision-language-models","repo_url":"https://github.com/muzairkhattak/multimodal-prompt-learning","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"learning-to-prompt-for-vision-language-models","repo_url":"https://github.com/muzairkhattak/protext","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"learning-to-prompt-for-vision-language-models","repo_url":"https://github.com/saic-fi/bayesian-prompt-learning","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"learning-to-prompt-for-vision-language-models","repo_url":"https://github.com/srvcodes/clap4clip","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}},{"paper_slug":"learning-to-prompt-for-vision-language-models","repo_url":"https://github.com/vill-lab/2024-aaai-hpt","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"domain-generalization","task_name":"Domain Generalization"},{"task_slug":"few-shot-age-estimation","task_name":"Few-shot Age Estimation"},{"task_slug":"prompt-engineering","task_name":"Prompt Engineering"},{"task_slug":"prompt-learning","task_name":"Prompt Learning"},{"task_slug":"representation-learning","task_name":"Representation Learning"}],"methods":[{"method_slug":"clip","method_name":"CLIP"},{"method_slug":"coop","method_name":"CoOp"}],"datasets_introduced":[],"methods_introduced":[{"slug":"coop","name":"CoOp","full_name":"Context Optimization"}],"results":[{"leaderboard":"/sota/few-shot-age-estimation-on-morph-album2","task":"Few-shot Age Estimation","dataset":"MORPH Album2","model":"CoOp","rank_in_archive_order":2,"of":2,"metrics":{"MAE":"5.09","MAE (16 shot)":"3.23","MAE (2 shot)":"4.50","MAE (4 shot)":"3.81","MAE (8 shot)":"3.57"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/2109.01134","atlas_url":"https://app.syntology.ai/?focus=2109.01134","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2109.01134"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-25T09:33:49+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/vill-lab/2024-aaai-hpt","reach":{"status":"ok","spdx":"MIT"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/kenomo/industrial-clip","reach":{"status":"ok","spdx":"MIT"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/hhenryd/tap","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/ArsenalCheng/Meta-Adapter","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/ThomasWangY/2024-AAAI-HPT","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/saic-fi/bayesian-prompt-learning","reach":{"status":"ok","spdx":"MIT"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/kaiyangzhou/coop","reach":{"status":"ok","spdx":"MIT"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/farinamatteo/zero","reach":{"status":"ok","spdx":"MIT"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/Gahyeonkim09/AAPL","reach":{"status":"ok","spdx":"MIT"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/kaiyangzhou/on-device-dg","reach":{"status":"ok","spdx":"NOASSERTION"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/srvcodes/clap4clip","reach":{"status":"ok"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/muzairkhattak/protext","reach":{"status":"ok","spdx":"MIT"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/healthx-lab/biomedcoop","reach":{"status":"ok","spdx":"MIT"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/YangYongJin/APEX","reach":{"status":"ok"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/mlvlab/dapt","reach":{"status":"ok","spdx":"MIT"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/Vill-Lab/2024-TIP-MetaPrompt","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/azshue/TPT","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/muzairkhattak/multimodal-prompt-learning","reach":{"status":"ok","spdx":"MIT"}}],"summary":{"ran":2,"ran_draft_wrong":1,"unverified":9},"by_repo_kind":{"listed":{"samples":12,"ran":3,"repositories":3}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"aa2102c5842a8045","entry":"CrossAttention","repo":"hhenryd/tap","repo_kind":"listed","path":"trainers/TAP.py","file_url":"https://github.com/hhenryd/tap/blob/HEAD/trainers/TAP.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":"invariant","behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"aa2102c5842a8045"}},{"code_sha256_prefix":"8a2b72f24481ca15","entry":"PromptLearner","repo":"ThomasWangY/2024-AAAI-HPT","repo_kind":"listed","path":"trainers/hpt.py","file_url":"https://github.com/ThomasWangY/2024-AAAI-HPT/blob/HEAD/trainers/hpt.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"8a2b72f24481ca15"}},{"code_sha256_prefix":"c568c664a705cfa6","entry":"make_description_batch","repo":"hhenryd/tap","repo_kind":"listed","path":"trainers/TAP.py","file_url":"https://github.com/hhenryd/tap/blob/HEAD/trainers/TAP.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"c568c664a705cfa6"}},{"code_sha256_prefix":"e8079fe48af584cb","entry":"AttentionPooling","repo":"hhenryd/tap","repo_kind":"listed","path":"trainers/TAP.py","file_url":"https://github.com/hhenryd/tap/blob/HEAD/trainers/TAP.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"e8079fe48af584cb"}},{"code_sha256_prefix":"7c02c9c890628418","entry":"CustomCLIP_TAP","repo":"hhenryd/tap","repo_kind":"listed","path":"trainers/TAP.py","file_url":"https://github.com/hhenryd/tap/blob/HEAD/trainers/TAP.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"7c02c9c890628418"}},{"code_sha256_prefix":"39be90830c4c1621","entry":"PromptLearner","repo":"Vill-Lab/2024-TIP-MetaPrompt","repo_kind":"listed","path":"trainers/meta.py","file_url":"https://github.com/Vill-Lab/2024-TIP-MetaPrompt/blob/HEAD/trainers/meta.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"39be90830c4c1621"}},{"code_sha256_prefix":"c1fabc9fbf753209","entry":"TextEncoder","repo":"hhenryd/tap","repo_kind":"listed","path":"trainers/TAP.py","file_url":"https://github.com/hhenryd/tap/blob/HEAD/trainers/TAP.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"c1fabc9fbf753209"}},{"code_sha256_prefix":"0746a2da2f73c0eb","entry":"TextPromptLearner","repo":"hhenryd/tap","repo_kind":"listed","path":"trainers/TAP.py","file_url":"https://github.com/hhenryd/tap/blob/HEAD/trainers/TAP.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"0746a2da2f73c0eb"}},{"code_sha256_prefix":"e1f5a7fb2fb95e19","entry":"VisualEncoder","repo":"hhenryd/tap","repo_kind":"listed","path":"trainers/TAP.py","file_url":"https://github.com/hhenryd/tap/blob/HEAD/trainers/TAP.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"e1f5a7fb2fb95e19"}},{"code_sha256_prefix":"d75235fc5acba1a6","entry":"VisualPromptLearner","repo":"hhenryd/tap","repo_kind":"listed","path":"trainers/TAP.py","file_url":"https://github.com/hhenryd/tap/blob/HEAD/trainers/TAP.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"d75235fc5acba1a6"}},{"code_sha256_prefix":"0f982cfc0f69d842","entry":"load_clip_to_cpu","repo":"hhenryd/tap","repo_kind":"listed","path":"trainers/TAP.py","file_url":"https://github.com/hhenryd/tap/blob/HEAD/trainers/TAP.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"0f982cfc0f69d842"}},{"code_sha256_prefix":"bf5154cf822c8f6a","entry":"load_gpt_descriptions","repo":"hhenryd/tap","repo_kind":"listed","path":"trainers/TAP.py","file_url":"https://github.com/hhenryd/tap/blob/HEAD/trainers/TAP.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"bf5154cf822c8f6a"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}