Datasets › ObjectNet

ObjectNet

Introduced by Andrei Barbu et al. in ObjectNet: A large-scale bias-controlled dataset for pushing the limits of object recognition models1 Jan 2019 archive 2025-07-28

ObjectNet is a test set of images collected directly using crowd-sourcing. ObjectNet is unique as the objects are captured at unusual poses in cluttered, natural scenes, which can severely degrade recognition performance. There are 50,000 images in the test set which controls for rotation, background and viewpoint. There are 313 object classes with 113 overlapping ImageNet.

Source: On Robustness and Transferability of Convolutional Neural Networks Image Source: https://objectnet.dev/

Benchmarks archive 2025-07-28

All 4 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

Papers archive 2025-07-28

30 shown of 41 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 155. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters 2 1 6 Feb 2024 ran 1 of 6 samples (5 unverified)
InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks 2 1 21 Dec 2023 ran 2 of 2 samples (0 unverified; 2 pointer-only for licence)
EVA-CLIP: Improved Training Techniques for CLIP at Scale 4 2 27 Mar 2023 ran 0 of 4 samples (4 unverified)
The effectiveness of MAE pre-pretraining for billion-scale pretraining 1 3 23 Mar 2023 not harvested
Scaling Vision Transformers to 22 Billion Parameters 1 1 10 Feb 2023 not harvested
A Whac-A-Mole Dilemma: Shortcuts Come in Multiples Where Mitigating One Amplifies Others 1 1 9 Dec 2022 not harvested
PaLI: A Jointly-Scaled Multilingual Language-Image Model 1 3 14 Sep 2022 ran 2 of 4 samples (2 unverified)
Optimizing Relevance Maps of Vision Transformers Improves Robustness 1 14 2 Jun 2022 not harvested
Matryoshka Representation Learning 5 1 26 May 2022 ran 2 of 3 samples (1 unverified)
CoCa: Contrastive Captioners are Image-Text Foundation Models 6 2 4 May 2022 ran 9 of 17 samples (8 unverified)
Data Determines Distributional Robustness in Contrastive Language Image Pre-training (CLIP) 2 1 3 May 2022 not harvested
Robust Cross-Modal Representation Learning with Progressive Self-Distillation 0 1 10 Apr 2022 not harvested
Representation Learning by Detecting Incorrect Location Embeddings 1 1 10 Apr 2022 not harvested
Bamboo: Building Mega-Scale Vision Dataset Continually with Human-Machine Synergy 2 2 15 Mar 2022 ran 3 of 3 samples (0 unverified; 3 pointer-only for licence)
Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time 6 2 10 Mar 2022 ran 5 of 17 samples (12 unverified)
Vision Models Are More Robust And Fair When Pretrained On Uncurated Images Without Supervision 1 1 16 Feb 2022 not harvested
Revisiting Weakly Supervised Pre-Training of Visual Perception Models 2 5 20 Jan 2022 not harvested
Pushing the limits of self-supervised ResNets: Can we outperform supervised learning without labels on ImageNet? 1 4 13 Jan 2022 ran 0 of 14 samples (14 unverified)
Optimal Representations for Covariate Shift 2 2 31 Dec 2021 ran 4 of 6 samples (2 unverified)
Pyramid Adversarial Training Improves ViT Performance 1 22 30 Nov 2021 not harvested
Discrete Representations Strengthen Vision Transformer Robustness 1 1 20 Nov 2021 not harvested
Combined Scaling for Zero-shot Transfer Learning 0 2 19 Nov 2021 not harvested
LiT: Zero-Shot Transfer with Locked-image text Tuning 5 2 15 Nov 2021 not harvested
Measuring the Interpretability of Unsupervised Representations via Quantized Reversed Probing 0 6 29 Sep 2021 not harvested
Compressive Visual Representations 1 2 27 Sep 2021 ran 0 of 12 samples (12 unverified)
Robust fine-tuning of zero-shot models 3 1 4 Sep 2021 ran 1 of 3 samples (2 unverified; 1 pointer-only for licence)
Billion-Scale Pretraining with Vision Transformers for Multi-Task Visual Representations 0 4 12 Aug 2021 not harvested
Compact and Optimal Deep Learning with Recurrent Parameter Generators 1 1 15 Jul 2021 ran 2 of 3 samples (1 unverified)
Scaling Vision Transformers 1 2 8 Jun 2021 not harvested
Characterizing and Improving the Robustness of Self-Supervised Learning through Background Augmentations 0 3 23 Mar 2021 not harvested

The full list of 41 is in the JSON twin.

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

License archive 2025-07-28

Creative Commons Attribution 4.0

Modalities archive 2025-07-28

Languages archive 2025-07-28

No language tagged.

Variants archive 2025-07-28

  • ObjectNet
  • ObjectNet (Bounding Box)

2 variant names, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections