Home › Datasets › task › Object Recognition

Object Recognition datasets

archive 2025-07-28

45 datasets carry the task tag "Object Recognition" (the task itself: Object Recognition), ordered by the archive's paper count. Page 1 of 1: 45 shown of 45. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 51 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets

Object Recognition datasets 1–45 of 45

description withheld: archive row vandalised before snapshot
16,145 papers · 91 benchmarks
The ImageNet dataset contains 14,197,122 annotated images according to the WordNet hierarchy.
15,430 papers · 52 benchmarks
CORe50 is a dataset designed for assessing Continual Learning techniques in an Object Recognition context.
132 papers · 0 benchmarks
The 'shape bias' dataset was introduced in Geirhos et al.
120 papers · 1 benchmark
N-Caltech 101 (Neuromorphic-Caltech101)
The Neuromorphic-Caltech101 (N-Caltech101) dataset is a spiking version of the original frame-based Caltech101 dataset.
110 papers · 3 benchmarks
Comprises 11 hand gesture categories from 29 subjects under 3 illumination conditions.
103 papers · 6 benchmarks
A large real-world event-based dataset for object classification.
56 papers · 2 benchmarks
The SUN Attribute dataset consists of 14,340 images from 717 scene categories, and each category is annotated with a taxonomy of 102 discriminate attributes.
38 papers · 2 benchmarks
OCID (Object Clutter Indoor Dataset)
Developing robot perception systems for handling objects in the real-world requires computer vision algorithms to be carefully scrutinized with respect to the expected operating domain.
29 papers · 1 benchmark
CIFAR10-DVS is an event-stream dataset for object classification.
23 papers · 2 benchmarks
Includes accurate pixel-wise motion masks, egomotion and ground truth depth.
22 papers · 0 benchmarks
The MECCANO dataset is the first dataset of egocentric videos to study human-object interactions in industrial-like settings.
19 papers · 3 benchmarks
N-ImageNet (Large-Scale Dataset for Event-Based Object Recognition)
The N-ImageNet dataset is an event-camera counterpart for the ImageNet dataset.
15 papers · 2 benchmarks
Washington RGB-D is a widely used testbed in the robotic community, consisting of 41,877 RGB-D images organized into 300 instances divided in 51 classes of common indoor objects (e.g.
15 papers · 0 benchmarks
DeepScores contains high quality images of musical scores, partitioned into 300,000 sheets of written music that contain symbols of different shapes and sizes.
10 papers · 0 benchmarks
TUM-GAID (TUM Gait from Audio, Image and Depth) collects 305 subjects performing two walking trajectories in an indoor environment.
9 papers · 0 benchmarks
APRICOT is a collection of over 1,000 annotated photographs of printed adversarial patches in public locations.
8 papers · 0 benchmarks
The RIT-18 dataset was built for the semantic segmentation of remote sensing imagery.
8 papers · 0 benchmarks
Freiburg Groceries is a groceries classification dataset consisting of 5000 images of size 256x256, divided into 25 categories.
7 papers · 0 benchmarks
iCubWorld datasets are collections of images recording the visual experience of iCub while observing objects in its typical environment, a laboratory or an office.
6 papers · 0 benchmarks
A large-scale video dataset for MOR in aerial videos.
6 papers · 0 benchmarks
ARID (Autonomous Robot Indoor Dataset)
ARID is a large-scale, multi-view object dataset collected with an RGB-D camera mounted on a mobile robot.
5 papers · 0 benchmarks
CURE-OR (Challenging Unreal and Real Environments for Object Recognition)
CURE-OR is a large-scale, controlled, and multi-platform object recognition dataset denoted as Challenging Unreal and Real Environments for Object Recognition.
5 papers · 0 benchmarks
The NYU Symmetry database contains 176 single-symmetry and 63 multiple-symmetry images (.png files) with accompanying ground-truth annotations (.mat files).
5 papers · 0 benchmarks
FewSOL (A Dataset for Few-Shot Object Learning in Robotic Environments)
The Few-Shot Object Learning (FewSOL) dataset can be used for object recognition with a few images per object.
4 papers · 0 benchmarks
HASY is a dataset of single symbols similar to MNIST.
3 papers · 0 benchmarks
A new large-scale retail product dataset for fine-grained image classification.
3 papers · 0 benchmarks
Egoshots is a 2-month Ego-vision Dataset with Autographer Wearable Camera annotated "for free" with transfer learning.
2 papers · 0 benchmarks
UW Indoor Scenes (UW-IS) Occluded dataset is curated using commodity hardware (Intel RealSense D435) to reflect real world robotics scenarios.
2 papers · 0 benchmarks
UW-IS (UW Indoor Scenes)
UW-IS (UW Indoor Scenes) is a dataset for object recognition in indoor environments comprising scene images from two different environments, namely, a living room and a mock warehouse.
2 papers · 0 benchmarks
The ARC-100 dataset was collected as part of a prototype retail checkout system titled ARC (Automatic Retail Checkout).
1 paper · 0 benchmarks
Multimodal object recognition is still an emerging field.
1 paper · 0 benchmarks
In this dataset two robots, Baxter and UR5, perform 8 behaviors (look, grasp, pick, hold, shake, lower, drop, and push) on 95 objects that vary by 5 color (blue, green, red, white, and yellow), 6 contents (wooden button, plastic dices,…
1 paper · 0 benchmarks
ChessReD (Chess Recognition Dataset)
The Chess Recognition Dataset (ChessReD) comprises a diverse collection of images of chess formations captured using smartphone cameras; a sensor choice made to ensure real-world applicability.
1 paper · 0 benchmarks
ChessReD2K (Chess Recognition Dataset 2K)
The Chess Recognition Dataset 2K (ChessReD2K) comprises a diverse collection of images of chess formations captured using smartphone cameras; a sensor choice made to ensure real-world applicability.
1 paper · 0 benchmarks
EGO-CH (EGOcentric-Cultural Heritage)
EGO-CH is a dataset of egocentric videos for visitors’ behavior understanding.
1 paper · 0 benchmarks
A new large-scale dataset that consists of 409 fine-grained categories and 31,881 images with accurate 3D pose annotation.
1 paper · 0 benchmarks
GOZ (Generic Object ZSL Dataset)
The Generix Object Zero-shot Learning (GOZ) dataset is a benchmark dataset for zero-shot learning.
1 paper · 0 benchmarks
Open MIC (Open Museum Identification Challenge)
Open MIC (Open Museum Identification Challenge) contains photos of exhibits captured in 10 distinct exhibition spaces of several museums which showcase paintings, timepieces, sculptures, glassware, relics, science exhibits, natural history…
1 paper · 0 benchmarks
Processed Twitter is a dataset that is used for Twitter topic recognition.
1 paper · 0 benchmarks
This dataset is used for RF signal recognition, used to recognize different RF devices based on the signals they transmitted.
1 paper · 0 benchmarks
STN PLAD (STN Power Line Assets Dataset)
STN PLAD is a high-resolution and real-world image dataset of multiple high-voltage power line components.
1 paper · 1 benchmark
Turath-150K is a database of images of the Arab world that reflect objects, activities, and scenarios commonly found there.
1 paper · 0 benchmarks
In this dataset UR5 robot used 6 tools: metal-scissor, metal-whisk, plastic-knife, plastic-spoon, wooden-chopstick, and wooden-fork to perform 6 behaviors: look, stirring-slow, stirring-fast, stirring-twist, whisk, and poke.
1 paper · 0 benchmarks
Dataset for testing the ability of Vision Language Models (LVM) to recognize and match 3D objects of the exact same 3D shapes but with different orientation/materials/textures/ environments and light conditions.
1 paper · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.