Browse State-of-the-Art › Zero-Shot Transfer Image Classification
Zero-Shot Transfer Image Classification
16 papers with code · 16 benchmarks · 8 datasets archive 2025-07-28
Benchmarks archive 2025-07-28
16 leaderboard tables shown for this task, 16 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted. 10 shown of 16 until expanded.
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
8 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Most implemented papers archive 2025-07-28
16 shown of 16 papers with code (19 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
26 Feb 2021 82 repositories listed Syntology ran 16 of 20 samples · 4 unverified · 16 pointer-only (licence)State-of-the-art computer vision systems are trained to predict a fixed set of predetermined object categories.
-
4 May 2022 6 repositories listed Syntology ran 9 of 17 samples · 8 unverifiedWe apply a contrastive loss between unimodal image and text embeddings, in addition to a captioning loss on the multimodal decoder outputs which predicts text tokens autoregressively.
-
15 Nov 2021 5 repositories listedThis paper presents contrastive-tuning, a simple method employing contrastive training to align image and text models while still taking advantage of their pre-training.
-
11 Feb 2021 5 repositories listed Syntology ran 8 of 10 samples · 2 unverified · 9 pointer-only (licence)In this paper, we leverage a noisy dataset of over one billion image alt-text pairs, obtained without expensive filtering or post-processing steps in the Conceptual Captions dataset.
-
28 Mar 2023 4 repositories listed Syntology ran 2 of 2 samples · 0 unverified · 1 pointer-only (licence)Our generative approach to classification, which we call Diffusion Classifier, attains strong results on a variety of benchmarks and outperforms alternative methods of extracting knowledge from diffusion models.
-
27 Mar 2023 4 repositories listed Syntology ran 0 of 4 samples · 4 unverifiedOur approach incorporates new techniques for representation learning, optimization, and augmentation, enabling EVA-CLIP to achieve superior performance compared to previous CLIP models with the same number of parameters…
-
6 Feb 2024 2 repositories listed Syntology ran 1 of 6 samples · 5 unverifiedScaling up contrastive language-image pretraining (CLIP) is critical for empowering both vision and multimodal models.
-
21 Dec 2023 2 repositories listed Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)However, the progress in vision and vision-language foundation models, which are also critical elements of multi-modal AGI, has not kept pace with LLMs.
-
12 Nov 2022 2 repositories listed Syntology ran 2 of 11 samples · 9 unverifiedIn this work, we present a conceptually simple and effective method to train a strong bilingual/multilingual multimodal representation model.
-
22 Nov 2021 2 repositories listedComputer vision foundation models, which are trained on diverse, large-scale dataset and can be adapted to a wide range of downstream tasks, are critical for this mission to solve real-world computer vision applications.
-
29 Jan 2024 1 repository listedVision-language foundation models like CLIP have revolutionized the field of artificial intelligence.
-
6 Jul 2023 1 repository listed Syntology ran 4 of 6 samples · 2 unverifiedModel distillation, the process of creating smaller, faster models that maintain the performance of larger models, is a promising direction towards the solution.
-
23 Mar 2023 1 repository listedWhile MAE has only been shown to scale with the size of models, we find that it scales with the size of the training dataset as well.
-
10 Feb 2023 1 repository listedThe scaling of Transformers has driven breakthrough capabilities for language models.
-
17 Jan 2023 1 repository listed Syntology ran 6 of 17 samples · 11 unverifiedImage-text contrastive learning models such as CLIP have demonstrated strong task transfer ability.
-
14 Sep 2022 1 repository listed Syntology ran 2 of 4 samples · 2 unverifiedPaLI generates text based on visual and textual inputs, and with this interface performs many vision, language, and multimodal tasks, in many languages.
Syntology lines on 11 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections