{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/context-aware-robust-fine-tuning","title":"Context-Aware Robust Fine-Tuning","arxiv_id":"2211.16175","date":"2022-11-29","proceeding":null,"authors":["Xiaofeng Mao","Yuefeng Chen","Xiaojun Jia","Rong Zhang","Hui Xue","Zhao Li"],"abstract":"Contrastive Language-Image Pre-trained (CLIP) models have zero-shot ability of classifying an image belonging to \"[CLASS]\" by using similarity between the image and the prompt sentence \"a [CONTEXT] of [CLASS]\". Based on exhaustive text cues in \"[CONTEXT]\", CLIP model is aware of different contexts, e.g. background, style, viewpoint, and exhibits unprecedented robustness against a wide range of distribution shifts. However, recent works find further fine-tuning of CLIP models improves accuracy but sacrifices the robustness on downstream tasks. We conduct an empirical investigation to show fine-tuning will corrupt the context-aware ability of pre-trained CLIP features. To solve this problem, we propose Context-Aware Robust Fine-tuning (CAR-FT). CAR-FT regularizes the model during fine-tuning to capture the context information. Specifically, we use zero-shot prompt weights to get the context distribution contained in the image. By minimizing the Kullback-Leibler Divergence (KLD) between context distributions induced by original/fine-tuned CLIP models, CAR-FT makes the context-aware ability of CLIP inherited into downstream tasks, and achieves both higher In-Distribution (ID) and Out-Of-Distribution (OOD) accuracy. The experimental results show CAR-FT achieves superior robustness on five OOD test datasets of ImageNet, and meanwhile brings accuracy gains on nine downstream tasks. Additionally, CAR-FT surpasses previous Domain Generalization (DG) methods and gets 78.5% averaged accuracy on DomainBed benchmark, building the new state-of-the-art.","url_abs":"https://arxiv.org/abs/2211.16175v1","url_pdf":"https://arxiv.org/pdf/2211.16175v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"domain-generalization","task_name":"Domain Generalization"},{"task_slug":"sentence","task_name":"Sentence"}],"methods":[{"method_slug":"aware","method_name":"AWARE"},{"method_slug":"clip","method_name":"CLIP"},{"method_slug":"test","method_name":"Test"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/domain-generalization-on-domainnet","task":"Domain Generalization","dataset":"DomainNet","model":"CAR-FT (CLIP, ViT-B/16)","rank_in_archive_order":6,"of":38,"metrics":{"Average Accuracy":"62.5"},"uses_additional_data":true},{"leaderboard":"/sota/domain-generalization-on-imagenet-a","task":"Domain Generalization","dataset":"ImageNet-A","model":"CAR-FT (CLIP, ViT-L/14@336px)","rank_in_archive_order":4,"of":39,"metrics":{"Top-1 accuracy %":"81.5"},"uses_additional_data":true},{"leaderboard":"/sota/domain-generalization-on-imagenet-r","task":"Domain Generalization","dataset":"ImageNet-R","model":"CAR-FT (CLIP, ViT-L/14@336px)","rank_in_archive_order":3,"of":39,"metrics":{"Top-1 Error Rate":"10.3"},"uses_additional_data":true},{"leaderboard":"/sota/domain-generalization-on-imagenet-sketch","task":"Domain Generalization","dataset":"ImageNet-Sketch","model":"CAR-FT (CLIP, ViT-L/14@336px)","rank_in_archive_order":3,"of":20,"metrics":{"Top-1 accuracy":"65.5"},"uses_additional_data":true},{"leaderboard":"/sota/domain-generalization-on-office-home","task":"Domain Generalization","dataset":"Office-Home","model":"CAR-FT (CLIP, ViT-B/16)","rank_in_archive_order":6,"of":45,"metrics":{"Average Accuracy":"85.7"},"uses_additional_data":false},{"leaderboard":"/sota/domain-generalization-on-pacs-2","task":"Domain Generalization","dataset":"PACS","model":"CAR-FT (CLIP, ViT-B/16)","rank_in_archive_order":10,"of":133,"metrics":{"Average Accuracy":"96.8"},"uses_additional_data":false},{"leaderboard":"/sota/domain-generalization-on-terraincognita","task":"Domain Generalization","dataset":"TerraIncognita","model":"CAR-FT (CLIP, ViT-B/16)","rank_in_archive_order":4,"of":30,"metrics":{"Average Accuracy":"61.9"},"uses_additional_data":false},{"leaderboard":"/sota/domain-generalization-on-vlcs","task":"Domain Generalization","dataset":"VLCS","model":"CAR-FT (CLIP, ViT-B/16)","rank_in_archive_order":1,"of":37,"metrics":{"Average Accuracy":"85.5"},"uses_additional_data":false}],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=2211.16175","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}