{"url":"/dataset/imagenet-w","name":"ImageNet-W","full_name":"ImageNet-Watermark","description_markdown":"ImageNet-W(atermark) is a test set to evaluate models’ reliance on the newly found watermark shortcut in ImageNet, which is used to predict the *carton* class. ImageNet-W is created by overlaying transparent watermarks on the ImageNet validation set. Two metrics are used to evaluate watermark shortcut reliance: (1) IN-W Gap: the top-1 accuracy drop from ImageNet to ImageNet-W, (2) Carton Gap: carton class accuracy increase from ImageNet to ImageNet-W. Combining ImageNet-W with previous out-of-distribution variants of ImageNet (e.g., Stylized ImageNet, ImageNet-R, ImageNet-9) forms a comprehensive suite of multi-shortcut evaluation on ImageNet.","description_withheld":null,"homepage":"https://github.com/facebookresearch/Whac-A-Mole","introduced_date":"2022-12-09","introduced_date_note":null,"introduced_by":{"paper":"/paper/a-whac-a-mole-dilemma-shortcuts-come-in","title":"A Whac-A-Mole Dilemma: Shortcuts Come in Multiples Where Mitigating One Amplifies Others","first_author":"Zhiheng Li","url":null},"license":{"name":"CC BY-NC","url":"https://github.com/facebookresearch/Whac-A-Mole/blob/main/LICENSE"},"modalities":[{"name":"Images","url":"/datasets/modality/images"}],"tasks":[{"name":"Image Classification","url":"/task/image-classification","datasets_with_task":"/datasets/task/image-classification"},{"name":"Out-of-Distribution Generalization","url":"/task/out-of-distribution-generalization","datasets_with_task":"/datasets/task/out-of-distribution-generalization"}],"languages":[],"variants":["ImageNet-W"],"data_loaders":[{"repo":"https://github.com/facebookresearch/Whac-A-Mole","url":"https://github.com/facebookresearch/Whac-A-Mole","frameworks":["pytorch"]}],"num_papers_in_archive":23,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}