Papers › Leveraging Vision-Language Models for Improving Domain Generalization in Image Classification

Leveraging Vision-Language Models for Improving Domain Generalization in Image Classification

12 Oct 2023CVPR 2024 1arXiv:2310.08255archive 2025-07-28

Sravanti Addepalli, Ashish Ramayee Asokan, Lakshay Sharma, R. Venkatesh Babu

Vision-Language Models (VLMs) such as CLIP are trained on large amounts of image-text pairs, resulting in remarkable generalization across several data distributions. However, in several cases, their expensive training and data collection/curation costs do not justify the end application. This motivates a vendor-client paradigm, where a vendor trains a large-scale VLM and grants only input-output access to clients on a pay-per-query basis in a black-box setting. The client aims to minimize inference cost by distilling the VLM to a student model using the limited available task-specific data, and further deploying this student model in the downstream application. While naive distillation largely improves the In-Domain (ID) accuracy of the student, it fails to transfer the superior out-of-distribution (OOD) generalization of the VLM teacher using the limited available labeled images. To mitigate this, we propose Vision-Language to Vision - Align, Distill, Predict (VL2V-ADiP), which first aligns the vision and language modalities of the teacher model with the vision modality of a pre-trained student model, and further distills the aligned VLM representations to the student. This maximally retains the pre-trained features of the student, while also incorporating the rich representations of the VLM image encoder and the superior generalization of the text embeddings. The proposed approach achieves state-of-the-art results on the standard Domain Generalization benchmarks in a black-box teacher setting as well as a white-box setting where the weights of the VLM are accessible.

PaperPDFConference PDFCodeCode Syntology ran

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

For agents, Syntology's MCP tool lists every function and class Syntology harvested from this paper and whether it ran (how to connect): get_harvested_code_for_paper(arxiv_id="2310.08255")

Code

Syntology Ran 12 of 15 code samples harvested from 1 repository linked to this paper; 3 have no recorded run. Of those that ran: 1 ran · honoured contract; 1 ran · our draft was wrong; 10 ran with no contract checked.

By repository: official repository: 15 samples from 1 repository, 12 ran. The run record, sample by sample. “Ran” means executed on a synthesized input, not that the code is correct or reproduces the paper.

val-iisc/VL2V-ADiP officialmentioned in papermentioned on GitHubpytorchMIT report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

15 samples harvested; 12 ran; 1 honoured the contract we drafted; 3 have no recorded run. Read from Syntology's graph 2026-09-24; that is when this build read the record, not when the samples ran.

1ran · honoured contract
1ran · our draft was wrong
10ran
3unverified

Licence: 0 of the 15 samples are pointer only, meaning Syntology does not serve that copy's text. This page shows no code text for any sample; each one links to its file in the repository.

Harvested from val-iisc/VL2V-ADiP. “Ran” means the sample executed on a synthesized input. It does not mean the output is correct, and nothing here reproduces the paper's results. “Honoured” and “violated” refer to a contract Syntology drafted from the code itself; “our draft was wrong” and “fixture could not drive it” are failures of Syntology's instrument, not of the code.

Each sample ends with its code_sha256, Syntology's identity for that exact code. An agent fetches the stored sample with Syntology's MCP tool get_code(code_sha256="…") (how to connect); click an identity to copy that call.

Repository labels, per sample. official repository: The archive marks this repository official for the paper. named in the paper: The archive records that the paper mentions this repository; it is not marked official. community (archive-listed): In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper. found in paper text by Syntology: Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted. community: Not in the archive's code links for this paper; a community repository Syntology harvested. Samples from a repository marked official are listed first. Licence labels name the repository's licence as recorded at harvest. “Pointer only” means Syntology does not serve that copy's text, for one of four reasons: no licence file was found; the licence was not identified; the licence is recorded as permissive but that copy's record is not marked cleared; or the licence is outside the permissive list Syntology serves text under (MIT, Apache-2.0, BSD and similar). Some licences outside that list permit redistribution, such as WTFPL, and GPL-3.0 under its conditions; they are simply not on the list. Hover a licence label for the reason. File links open the file on GitHub at the default branch, which may have changed since the harvest.

default_hparams val-iisc/VL2V-ADiP/domainbed/hparams_registry.py official repository ran fingerprinted MIT (permissive) · d6339c46586cd5f7 · report
get_shapes val-iisc/VL2V-ADiP/domainbed/algorithms/miro.py official repository ran MIT (permissive) · c5919b96db289dbd · report
hashable val-iisc/VL2V-ADiP/domainbed/lib/query.py official repository ran fingerprinted MIT (permissive) · 71a3a61ceed99bf3 · report
levelize val-iisc/VL2V-ADiP/domainbed/lib/logger.py official repository ran fingerprinted MIT (permissive) · 2f597a2077a1c95d · report
make_selector_fn val-iisc/VL2V-ADiP/domainbed/lib/query.py official repository ran MIT (permissive) · 2f8af5779e1edcb6 · report
make_weights_for_balanced_classes val-iisc/VL2V-ADiP/domainbed/lib/misc.py official repository ran MIT (permissive) · 955c4020424d8ac1 · report
random_hparams val-iisc/VL2V-ADiP/domainbed/hparams_registry.py official repository ran MIT (permissive) · 72bf0287a79e4129 · report
random_hparams_st2 val-iisc/VL2V-ADiP/domainbed/hparams_registry.py official repository ran MIT (permissive) · 185f6c7403da6c33 · report
random_pairs_of_minibatches val-iisc/VL2V-ADiP/domainbed/lib/misc.py official repository ran · our draft was wrong MIT (permissive) · 4fa2b54178fcc1df · report
str2bool val-iisc/VL2V-ADiP/train_all.py official repository ran MIT (permissive) · a9fcb7b2aef70def · report
to_minibatch val-iisc/VL2V-ADiP/domainbed/algorithms/algorithms.py official repository ran · honoured contract fingerprinted MIT (permissive) · 12c23ce22097acd0 · report
to_row val-iisc/VL2V-ADiP/domainbed/lib/misc.py official repository ran MIT (permissive) · 5dd6a674fbe45a89 · report
accuracy_from_loader val-iisc/VL2V-ADiP/domainbed/evaluator.py official repository unverified MIT (permissive) · 687efdf3db9c44df · report
accuracy_from_loader_cls val-iisc/VL2V-ADiP/domainbed/evaluator.py official repository unverified MIT (permissive) · a5527e8f11826a39 · report
one_hot_embedding val-iisc/VL2V-ADiP/domainbed/algorithms/baselines.py official repository unverified MIT (permissive) · 7c5a3128571afe6f · report

Tasks

Domain GeneralizationImage Classificationimage-classification

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Domain Generalization DomainNet VL2V-SD (CLIP, ViT-B/16) Average Accuracy 62.79 #4 of 38 Archive leaderboard report
Domain Generalization Office-Home VL2V-SD (CLIP, ViT-B/16) Average Accuracy 87.38 #5 of 45 Archive leaderboard report
Domain Generalization PACS VL2V-SD (CLIP, ViT-B/16) Average Accuracy 96.68 #11 of 133 Archive leaderboard report
Domain Generalization TerraIncognita VL2V-SD (CLIP, ViT-B/16) Average Accuracy 58.54 #8 of 30 Archive leaderboard report
Domain Generalization VLCS VL2V-SD (CLIP, ViT-B/16) Average Accuracy 83.25 #4 of 37 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

CLIP

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections