{"url":"/dataset/imagenet-atr","name":"ImageNet-Atr","full_name":"ImageNet with Adversarial Text Regions","description_markdown":"We build a new evaluation set by adding spotting words to the images of ImageNet 2012 evaluation sets. There are 1,000 categories in ImageNet. For each category c, we find its most confusing category c*and spot the category name to every evaluation image. \r\n\r\nThis evaluation set is challenging for many CLIP models. For example, OpenAI CLIP B-16 got a top-1 accuracy of as low as 32%, which is much lower than the original ImageNet evaluation set.","description_withheld":null,"homepage":"https://github.com/apple/axlearn/tree/main/axlearn/vision/imagenet_adversarial_text","introduced_date":"2023-05-08","introduced_date_note":null,"introduced_by":{"paper":"/paper/less-is-more-removing-text-regions-improves","title":"Less is More: Removing Text-regions Improves CLIP Training Efficiency and Robustness","first_author":"Liangliang Cao","url":null},"license":null,"modalities":[{"name":"Images","url":"/datasets/modality/images"}],"tasks":[],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["ImageNet-Atr"],"data_loaders":[],"num_papers_in_archive":1,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-25T09:33:49+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}