Browse State-of-the-Art › Object Recognition
Object Recognition
577 papers with code · 9 benchmarks · 45 datasets archive 2025-07-28
Object recognition is a computer vision technique for detecting + classifying objects in images or videos. Since this is a combined task of object detection plus image classification, the state-of-the-art tables are recorded for each component task here and here.
( Image credit: Tensorflow Object Detection API )
Description from the archive archive 2025-07-28; Papers-with-Code links inside it are rewritten to this site.
Benchmarks archive 2025-07-28
9 leaderboard tables shown for this task, 9 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| shape bias (18 rows) | Imagen | Intriguing properties of generative classifiers | code | Syntology ran 10 of 12 samples · 2 unverified | Compare |
| CIFAR10-DVS (2 rows) | Spike-VGG11 | EventRPG: Event Data Augmentation with Relevance Propagation Guidance | code | Syntology ran 5 of 16 samples · 11 unverified | Compare |
| N-Caltech 101 (2 rows) | Spike-VGG11 | EventRPG: Event Data Augmentation with Relevance Propagation Guidance | code | Syntology ran 5 of 16 samples · 11 unverified | Compare |
| ObjectNet (All classes) (2 rows) | ObjectNet-Baseline | — | — | — | Compare |
| ObjectNet (ImageNet classes, trained on ImageNet) (2 rows) | ObjectNet-Baseline | — | — | — | Compare |
| ObjectNet (ImageNet classes) (2 rows) | ObjectNet-Baseline | — | — | — | Compare |
| DVS128 Gesture (1 row) | SSNN | Shrinking Your TimeStep: Towards Low-Latency Neuromorphic Object... | — | — | Compare |
| MECCANO (1 row) | Faster-RCNN | The MECCANO Dataset: Understanding Human-Object Interactions from... | code | — | Compare |
| N-CARS (1 row) | Spike-VGG11 | EventRPG: Event Data Augmentation with Relevance Propagation Guidance | code | Syntology ran 5 of 16 samples · 11 unverified | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
45 datasets whose archive record lists this task, ordered by the archive's paper count. 30 shown of 45 until expanded.
Subtasks archive 2025-07-28
3 subtasks in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 577 papers with code (2,042 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
25 Aug 2016 146 repositories listed Syntology ran 18 of 71 samples · 53 unverified · 7 pointer-only (licence)Recent work has shown that convolutional networks can be substantially deeper, more accurate, and efficient to train if they contain shorter connections between layers close to the input and those close to the output.
-
13 Feb 2020 96 repositories listed Syntology ran 79 of 137 samples · 58 unverified · 52 pointer-only (licence)This paper presents SimCLR: a simple framework for contrastive learning of visual representations.
-
17 Sep 2014 83 repositories listed Syntology ran 27 of 42 samples · 15 unverified · 20 pointer-only (licence)We propose a deep convolutional neural network architecture codenamed "Inception", which was responsible for setting the new state of the art for classification and detection in the ImageNet Large-Scale Visual…
-
26 Feb 2021 82 repositories listed Syntology ran 16 of 20 samples · 4 unverified · 16 pointer-only (licence)State-of-the-art computer vision systems are trained to predict a fixed set of predetermined object categories.
-
1 May 2014 38 repositories listed Syntology ran 1 of 7 samples · 6 unverifiedWe present a new dataset with the goal of advancing the state-of-the-art in object recognition by placing the question of object recognition in the context of the broader question of scene understanding.
-
21 Dec 2014 37 repositories listed Syntology ran 0 of 6 samples · 6 unverifiedMost modern convolutional neural networks (CNNs) used for object recognition are built using the same principles: Alternating convolution and max-pooling layers followed by a small number of fully connected layers.
-
1 Dec 2012 23 repositories listedWe trained a large, deep convolutional neural network to classify the 1.
-
13 Dec 2016 20 repositories listed Syntology ran 0 of 14 samples · 14 unverifiedWe explore three aspects of the problem in the context of finding small faces: the role of scale invariance, image resolution, and contextual reasoning.
-
23 Apr 2017 19 repositories listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)In this work, we propose "Residual Attention Network", a convolutional neural network using attention mechanism which can incorporate with state-of-art feed forward network architecture in an end-to-end training fashion.
-
25 May 2016 17 repositories listed Syntology ran 0 of 6 samples · 6 unverified · 2 pointer-only (licence)Here, we explore prediction of future frames in a video sequence as an unsupervised learning rule for learning about the structure of the visual world.
-
27 May 2015 16 repositories listed Syntology ran 2 of 3 samples · 1 unverified · 3 pointer-only (licence)Here we introduce a new model of natural textures based on the feature spaces of convolutional neural networks optimised for object recognition.
-
14 Dec 2022 14 repositories listed Syntology ran 3 of 20 samples · 17 unverifiedIn this paper, we aim to design an efficient real-time object detector that exceeds the YOLO series and is easily extensible for many object recognition tasks such as instance segmentation and rotated object detection.
-
4 Dec 2018 14 repositories listed Syntology ran 7 of 8 samples · 1 unverified · 1 pointer-only (licence)Dot-product attention has wide applications in computer vision and natural language processing.
-
1 Sep 2014 14 repositories listed Syntology ran 1 of 5 samples · 4 unverifiedThe ImageNet Large Scale Visual Recognition Challenge is a benchmark in object category classification and detection on hundreds of object categories and millions of images.
-
18 Jun 2014 14 repositories listed Syntology ran 0 of 1 samples · 1 unverifiedThis requirement is "artificial" and may reduce the recognition accuracy for the images or sub-images of an arbitrary size/scale.
-
14 Nov 2013 14 repositories listedPatterns and textures are defining characteristics of many natural objects: a shirt can be striped, the wings of a butterfly can be veined, and the skin of an animal can be scaly.
-
3 Jul 2012 11 repositories listedWhen a large feedforward neural network is trained on a small training set, it typically performs poorly on held-out test data.
-
25 Apr 2019 9 repositories listedIn this paper, we take advantage of this finding to create a simplified network based on a query-independent formulation, which maintains the accuracy of NLNet but with significantly less computation.
-
3 Feb 2015 9 repositories listedVery deep neural networks recently achieved great success on general object recognition because of their superb learning capacity.
-
6 Oct 2013 8 repositories listed Syntology ran 1 of 7 samples · 6 unverified · 1 pointer-only (licence)We evaluate whether features extracted from the activation of a deep convolutional network trained in a fully supervised fashion on a large, fixed set of object recognition tasks can be re-purposed to novel generic…
-
29 Nov 2018 7 repositories listed Syntology ran 0 of 6 samples · 6 unverifiedConvolutional Neural Networks (CNNs) are commonly thought to recognise objects by learning increasingly complex representations of object shapes.
-
13 Aug 2017 7 repositories listedHowever, its solution is crucial for many experienced players who wish to compete against AI bots, but also prefer to make decisions based on the analysis of a physical chessboard.
-
25 Nov 2020 6 repositories listedIn our method, however, a fixed sparse set of learned object proposals, total length of N, are provided to object recognition head to perform classification and location.
-
20 Mar 2020 6 repositories listedTherefore, our multi-branch and multi-scale learning network(MMAL-Net) has good classification ability and robustness for images of different scales.
-
30 Nov 2017 6 repositories listedAlthough it is well believed for years that modeling relations between objects would help object recognition, there has not been evidence that the idea is working in the deep learning era.
-
8 Dec 2013 6 repositories listedDeep convolutional neural networks have recently achieved state-of-the-art performance on a number of image recognition benchmarks, including the ImageNet Large-Scale Visual Recognition Challenge (ILSVRC-2012).
-
8 Jul 2019 5 repositories listedIdeally, continual learning should be triggered by the availability of short videos of single objects and performed on-line on on-board hardware with fine-grained updates.
-
11 Apr 2017 5 repositories listedWe present an exhaustive investigation of recent Deep Learning architectures, algorithms, and strategies for the task of document image classification to finally reduce the error by more than half.
-
6 Aug 2016 5 repositories listedTo remedy this, we develop a method for adapting deep features to align with human similarity judgments, resulting in image representations that can potentially be used to extend the scope of psychological experiments.
-
24 Dec 2014 5 repositories listed Syntology ran 3 of 3 samples · 0 unverified · 3 pointer-only (licence)We present an attention-based model for recognizing multiple objects in images.
Syntology lines on 17 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections