{"url":"/dataset/vtab","name":"VTAB","full_name":"Visual Task Adaptation Benchmark","description_markdown":"The **Visual Task Adaptation Benchmark (VTAB)** is a benchmark designed to evaluate general visual representations². It consists of a diverse and challenging suite of tasks². The benchmark defines a good general visual representation as one that yields good performance on unseen tasks, when trained on limited task-specific data².\r\n\r\nThe VTAB benchmark contains the following 19 tasks that are derived from public datasets¹:\r\n- Caltech101\r\n- CIFAR-100\r\n- CLEVR distance prediction\r\n- CLEVR counting\r\n- Diabetic Rethinopathy\r\n- Dmlab Frames\r\n- dSprites orientation prediction\r\n- dSprites location prediction\r\n- Describable Textures Dataset (DTD)\r\n- EuroSAT\r\n- KITTI distance prediction\r\n- 102 Category Flower Dataset\r\n- Oxford IIIT Pet dataset\r\n- PatchCamelyon\r\n- Resisc45\r\n- Small NORB azimuth prediction\r\n- Small NORB elevation prediction\r\n- SUN397\r\n- SVHN\r\n\r\nThe given model is independently fine-tuned for solving each of the above tasks¹. Average accuracy across all tasks is used to measure the model's performance¹. Detailed description of all tasks, evaluation protocol, and other details can be found in the VTAB paper¹.\r\n\r\n(1) Visual Task Adaptation Benchmark. https://google-research.github.io/task_adaptation/.\r\n(2) GitHub - google-research/task_adaptation. https://github.com/google-research/task_adaptation.\r\n(3) GitHub - KMnP/vpt: ️ Visual Prompt Tuning [ECCV 2022] https://arxiv .... https://github.com/KMnP/vpt.","description_withheld":null,"homepage":"https://google-research.github.io/task_adaptation","introduced_date":"2019-09-25","introduced_date_note":null,"introduced_by":{"paper":"/paper/the-visual-task-adaptation-benchmark-1","title":"The Visual Task Adaptation Benchmark","first_author":"Xiaohua Zhai","url":null},"license":null,"modalities":[],"tasks":[{"name":"Image Classification","url":"/task/image-classification","datasets_with_task":"/datasets/task/image-classification"},{"name":"Visual Prompt Tuning","url":"/task/visual-prompt-tuning","datasets_with_task":"/datasets/task/visual-prompt-tuning"}],"languages":[],"variants":["VTAB-1k","VTAB-1k(Natural<7>)","VTAB-1k(Specialized<4>)","VTAB-1k(Structured<8>)","VTAB"],"data_loaders":[],"num_papers_in_archive":202,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/image-classification-on-vtab-1k-1","task":"Image Classification","dataset_variant":"VTAB-1k","rows":34,"metrics":["Top-1 Accuracy"],"first_row_in_archive_order":{"model":"ALIGN (50 hypers/task)","paper":"/paper/scaling-up-visual-and-vision-language","metrics":{"Top-1 Accuracy":"79.99"},"code_links":[{"title":"facebookresearch/metaclip","url":"https://github.com/facebookresearch/metaclip"},{"title":"kakaobrain/coyo-dataset","url":"https://github.com/kakaobrain/coyo-dataset"},{"title":"MicPie/clasp","url":"https://github.com/MicPie/clasp"},{"title":"willard-yuan/video-text-retrieval-papers","url":"https://github.com/willard-yuan/video-text-retrieval-papers"},{"title":"pwc-1/Paper-8","url":"https://github.com/pwc-1/Paper-8/tree/main/align"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/visual-prompt-tuning-on-vtab-1k-natural-7","task":"Visual Prompt Tuning","dataset_variant":"VTAB-1k(Natural<7>)","rows":10,"metrics":["Mean Accuracy"],"first_row_in_archive_order":{"model":"SPT-Deep(ViT-B/16_MoCo_v3_pretrained_ImageNet-1K)","paper":"/paper/revisiting-the-power-of-prompt-for-visual","metrics":{"Mean Accuracy":"76.20"},"code_links":[]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/visual-prompt-tuning-on-vtab-1k-specialized-4","task":"Visual Prompt Tuning","dataset_variant":"VTAB-1k(Specialized<4>)","rows":10,"metrics":["Mean Accuracy"],"first_row_in_archive_order":{"model":"SPT-Deep(ViT-B/16_MoCo_v3_pretrained_ImageNet-1K)","paper":"/paper/revisiting-the-power-of-prompt-for-visual","metrics":{"Mean Accuracy":"84.95"},"code_links":[]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/visual-prompt-tuning-on-vtab-1k-structured-8","task":"Visual Prompt Tuning","dataset_variant":"VTAB-1k(Structured<8>)","rows":10,"metrics":["Mean Accuracy"],"first_row_in_archive_order":{"model":"SPT-Deep(ViT-B/16_MAE_pretrained_ImageNet-1K)","paper":"/paper/revisiting-the-power-of-prompt-for-visual","metrics":{"Mean Accuracy":"59.23"},"code_links":[]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[{"paper":"/paper/revisiting-the-power-of-prompt-for-visual","title":"Revisiting the Power of Prompt for Visual Tuning","date":"2024-02-04","rows_on_this_dataset":12,"code_links":0,"syntology":null},{"paper":"/paper/improving-visual-prompt-tuning-for-self","title":"Improving Visual Prompt Tuning for Self-supervised Vision Transformers","date":"2023-06-08","rows_on_this_dataset":6,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":4,"samples_ran":4,"samples_unverified":0,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/pvp-pre-trained-visual-parameter-efficient","title":"PVP: Pre-trained Visual Parameter-Efficient Tuning","date":"2023-04-26","rows_on_this_dataset":1,"code_links":0,"syntology":null},{"paper":"/paper/contextual-squeeze-and-excitation-for","title":"Contextual Squeeze-and-Excitation for Efficient Few-Shot Image Classification","date":"2022-06-20","rows_on_this_dataset":2,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":3,"samples_ran":3,"samples_unverified":0,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/visual-prompt-tuning","title":"Visual Prompt Tuning","date":"2022-03-23","rows_on_this_dataset":12,"code_links":6,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":27,"samples_ran":17,"samples_unverified":10,"pointer_only_for_licence":15,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/scaling-vision-transformers","title":"Scaling Vision Transformers","date":"2021-06-08","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/scaling-up-visual-and-vision-language","title":"Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision","date":"2021-02-11","rows_on_this_dataset":1,"code_links":5,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":10,"samples_ran":8,"samples_unverified":2,"pointer_only_for_licence":9,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/deep-ensembles-for-low-data-transfer-learning-1","title":"Deep Ensembles for Low-Data Transfer Learning","date":"2020-10-14","rows_on_this_dataset":1,"code_links":0,"syntology":null},{"paper":"/paper/scalable-transfer-learning-with-expert-models","title":"Scalable Transfer Learning with Expert Models","date":"2020-09-28","rows_on_this_dataset":1,"code_links":0,"syntology":null},{"paper":"/paper/large-scale-learning-of-general-visual","title":"Big Transfer (BiT): General Visual Representation Learning","date":"2019-12-24","rows_on_this_dataset":4,"code_links":9,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":10,"samples_ran":3,"samples_unverified":7,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/self-supervised-learning-of-video-induced","title":"Self-Supervised Learning of Video-Induced Visual Invariances","date":"2019-12-05","rows_on_this_dataset":1,"code_links":0,"syntology":null},{"paper":"/paper/the-visual-task-adaptation-benchmark","title":"A Large-scale Study of Representation Learning with the Visual Task Adaptation Benchmark","date":"2019-10-01","rows_on_this_dataset":22,"code_links":2,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":1,"samples_ran":1,"samples_unverified":0,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}}],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":6,"samples_harvested":55,"samples_ran":36,"samples_unverified":19,"pointer_only_for_licence":24,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}