Datasets › VTAB
VTAB (Visual Task Adaptation Benchmark)
The Visual Task Adaptation Benchmark (VTAB) is a benchmark designed to evaluate general visual representations². It consists of a diverse and challenging suite of tasks². The benchmark defines a good general visual representation as one that yields good performance on unseen tasks, when trained on limited task-specific data².
The VTAB benchmark contains the following 19 tasks that are derived from public datasets¹: - Caltech101 - CIFAR-100 - CLEVR distance prediction - CLEVR counting - Diabetic Rethinopathy - Dmlab Frames - dSprites orientation prediction - dSprites location prediction - Describable Textures Dataset (DTD) - EuroSAT - KITTI distance prediction - 102 Category Flower Dataset - Oxford IIIT Pet dataset - PatchCamelyon - Resisc45 - Small NORB azimuth prediction - Small NORB elevation prediction - SUN397 - SVHN
The given model is independently fine-tuned for solving each of the above tasks¹. Average accuracy across all tasks is used to measure the model's performance¹. Detailed description of all tasks, evaluation protocol, and other details can be found in the VTAB paper¹.
(1) Visual Task Adaptation Benchmark. https://google-research.github.io/task_adaptation/. (2) GitHub - google-research/task_adaptation. https://github.com/google-research/task_adaptation. (3) GitHub - KMnP/vpt: ️ Visual Prompt Tuning [ECCV 2022] https://arxiv .... https://github.com/KMnP/vpt.
Benchmarks archive 2025-07-28
All 4 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.
| First row (archive order) | Paper | Code | ||||
|---|---|---|---|---|---|---|
| Image Classification | VTAB-1k | ALIGN (50 hypers/task) Top-1 Accuracy 79.99 | Scaling Up Visual and Vision-Language Representation... | facebookresearch/metaclip +4 | 34 | Compare |
| Visual Prompt Tuning | VTAB-1k(Natural<7>) | SPT-Deep(ViT-B/16_MoCo_v3_pretrained_ImageNet-1K) Mean Accuracy 76.20 | Revisiting the Power of Prompt for Visual Tuning | — | 10 | Compare |
| Visual Prompt Tuning | VTAB-1k(Specialized<4>) | SPT-Deep(ViT-B/16_MoCo_v3_pretrained_ImageNet-1K) Mean Accuracy 84.95 | Revisiting the Power of Prompt for Visual Tuning | — | 10 | Compare |
| Visual Prompt Tuning | VTAB-1k(Structured<8>) | SPT-Deep(ViT-B/16_MAE_pretrained_ImageNet-1K) Mean Accuracy 59.23 | Revisiting the Power of Prompt for Visual Tuning | — | 10 | Compare |
Papers archive 2025-07-28
12 shown of 12 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 202. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.
Dataset loaders archive 2025-07-28
No loader listed in the archive.
Tasks archive 2025-07-28
License archive 2025-07-28
No licence recorded in the archive. Absence here is not a statement about the dataset's terms.
Modalities archive 2025-07-28
No modality tagged.
Languages archive 2025-07-28
No language tagged.
Variants archive 2025-07-28
- VTAB-1k
- VTAB-1k(Natural<7>)
- VTAB-1k(Specialized<4>)
- VTAB-1k(Structured<8>)
- VTAB
5 variant names, as the archive lists them.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections