Datasets › ImageNet-Sketch

ImageNet-Sketch

Introduced by Haohan Wang et al. in Learning Robust Global Representations by Penalizing Local Predictive Power29 May 2019 archive 2025-07-28

ImageNet-Sketch data set consists of 50,889 images, approximately 50 images for each of the 1000 ImageNet classes. The data set is constructed with Google Image queries "sketch of __", where __ is the standard class name. Only within the "black and white" color scheme is searched. 100 images are initially queried for every class, and the pulled images are cleaned by deleting the irrelevant images and images that are for similar but different classes. For some classes, there are less than 50 images after manually cleaning, and then the data set is augmented by flipping and rotating the images.

Source: ImageNet-Sketch Image Source: https://github.com/HaohanWang/ImageNet-Sketch

Benchmarks archive 2025-07-28

All 3 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Domain Generalization ImageNet-Sketch Model soups (BASIC-L) Top-1 accuracy 77.18 Model soups: averaging weights of multiple fine-tuned... mlfoundations/model-soups +5 20 Compare
Zero-Shot Transfer Image Classification ImageNet-Sketch CoCa Accuracy (Private) 77.6 CoCa: Contrastive Captioners are Image-Text Foundation Models mlfoundations/open_clip +5 7 Compare
Image Classification ImageNet-Sketch µ2Net+ (ViT-L/16) Accuracy 88.6 A Continual Development Methodology for Large-scale... google-research/google-research 1 Compare

Papers archive 2025-07-28

20 shown of 20 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 268. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters 2 1 6 Feb 2024 ran 1 of 6 samples (5 unverified)
InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks 2 1 21 Dec 2023 ran 2 of 2 samples (0 unverified; 2 pointer-only for licence)
Distilling Out-of-Distribution Robustness from Vision-Language Foundation Models 1 1 2 Nov 2023 ran 1 of 1 samples (0 unverified)
EVA-CLIP: Improved Training Techniques for CLIP at Scale 4 1 27 Mar 2023 ran 0 of 4 samples (4 unverified)
A Whac-A-Mole Dilemma: Shortcuts Come in Multiples Where Mitigating One Amplifies Others 1 1 9 Dec 2022 not harvested
Context-Aware Robust Fine-Tuning 0 1 29 Nov 2022 not harvested
AltCLIP: Altering the Language Encoder in CLIP for Extended Language Capabilities 2 1 12 Nov 2022 ran 2 of 11 samples (9 unverified)
MetaFormer Baselines for Vision 8 6 24 Oct 2022 ran 0 of 4 samples (4 unverified)
Generalized Parametric Contrastive Learning 4 1 26 Sep 2022 not harvested
Enhance the Visual Representation via Discrete Adversarial Training 1 1 16 Sep 2022 ran 10 of 11 samples (1 unverified; 1 pointer-only for licence)
A Continual Development Methodology for Large-scale Multitask Dynamic ML Systems 1 1 15 Sep 2022 not harvested
Sequencer: Deep LSTM for Image Classification 5 1 4 May 2022 ran 4 of 9 samples (5 unverified)
CoCa: Contrastive Captioners are Image-Text Foundation Models 6 1 4 May 2022 ran 9 of 17 samples (8 unverified)
Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time 6 2 10 Mar 2022 ran 5 of 17 samples (12 unverified)
Vision Models Are More Robust And Fair When Pretrained On Uncurated Images Without Supervision 1 1 16 Feb 2022 not harvested
A ConvNet for the 2020s 54 1 10 Jan 2022 ran 54 of 80 samples (26 unverified; 11 pointer-only for licence)
Pyramid Adversarial Training Improves ViT Performance 1 2 30 Nov 2021 not harvested
Discrete Representations Strengthen Vision Transformer Robustness 1 1 20 Nov 2021 not harvested
Combined Scaling for Zero-shot Transfer Learning 0 1 19 Nov 2021 not harvested
Masked Autoencoders Are Scalable Vision Learners 58 1 11 Nov 2021 ran 71 of 137 samples (66 unverified; 73 pointer-only for licence)

Dataset loaders archive 2025-07-28

4 loaders as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

Unknown

Modalities archive 2025-07-28

Languages archive 2025-07-28

No language tagged.

Variants archive 2025-07-28

  • ImageNet-Sketch

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections