Browse State-of-the-Art › Object Detection
Object Detection
4,657 papers with code · 123 benchmarks · 332 datasets archive 2025-07-28
Benchmarks archive 2025-07-28
124 leaderboard tables shown for this task (3 more in the archive withheld as spam; see /not-shown), 123 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted. 10 shown of 124 until expanded.
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
332 datasets whose archive record lists this task, ordered by the archive's paper count. 30 shown of 332 until expanded.
Subtasks archive 2025-07-28
39 subtasks in the archive's task tree. 30 shown of 39 until expanded.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
30 shown of 4,657 papers with code (10,957 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
10 Dec 2015 484 repositories listed Syntology ran 230 of 377 samples · 147 unverified · 187 pointer-only (licence)Deep residual nets are foundations of our submissions to ILSVRC & COCO 2015 competitions, where we also won the 1st places on the tasks of ImageNet detection, ImageNet localization, COCO detection, and COCO segmentation.
-
8 Apr 2018 311 repositories listed Syntology ran 18 of 124 samples · 106 unverified · 19 pointer-only (licence)At 320x320 YOLOv3 runs in 22 ms at 28.
-
7 Aug 2017 234 repositories listed Syntology ran 11 of 11 samples · 0 unverified · 6 pointer-only (licence)Our novel Focal Loss focuses training on a sparse set of hard examples and prevents the vast number of easy negatives from overwhelming the detector during training.
-
25 Dec 2016 231 repositories listed Syntology ran 16 of 60 samples · 44 unverified · 22 pointer-only (licence)On the 156 classes not in COCO, YOLO9000 gets 16.
-
23 Apr 2020 223 repositories listed Syntology ran 24 of 184 samples · 160 unverified · 8 pointer-only (licence)There are a huge number of features which are said to improve Convolutional Neural Network (CNN) accuracy.
-
8 Dec 2015 221 repositories listed Syntology ran 19 of 131 samples · 112 unverified · 5 pointer-only (licence)Experimental results on the PASCAL VOC, MS COCO, and ILSVRC datasets confirm that SSD has comparable accuracy to methods that utilize an additional object proposal step and is much faster, while providing a unified…
-
4 Jun 2015 196 repositories listed Syntology ran 59 of 124 samples · 65 unverified · 42 pointer-only (licence)In this work, we introduce a Region Proposal Network (RPN) that shares full-image convolutional features with the detection network, thus enabling nearly cost-free region proposals.
-
20 Mar 2017 179 repositories listed Syntology ran 42 of 140 samples · 98 unverified · 23 pointer-only (licence)Our approach efficiently detects objects in an image while simultaneously generating a high-quality segmentation mask for each instance.
-
13 Jan 2018 159 repositories listed Syntology ran 85 of 111 samples · 26 unverified · 64 pointer-only (licence)In this paper we describe a new mobile architecture, MobileNetV2, that improves the state of the art performance of mobile models on multiple tasks and benchmarks as well as across a spectrum of different model sizes.
-
17 Apr 2017 159 repositories listed Syntology ran 52 of 83 samples · 31 unverified · 48 pointer-only (licence)We present a class of efficient models called MobileNets for mobile and embedded vision applications.
-
8 Jun 2015 144 repositories listed Syntology ran 80 of 148 samples · 68 unverified · 98 pointer-only (licence)A single neural network predicts bounding boxes and class probabilities directly from full images in one evaluation.
-
17 Jun 2019 142 repositories listed Syntology ran 14 of 82 samples · 68 unverifiedIn this paper, we introduce the various features of this toolbox.
-
27 Nov 2019 123 repositories listedNeural networks have enabled state-of-the-art approaches to achieve incredible results on computer vision tasks such as object detection.
-
2 Apr 2019 87 repositories listed Syntology ran 13 of 40 samples · 27 unverified · 18 pointer-only (licence)By eliminating the predefined set of anchor boxes, FCOS completely avoids the complicated computation related to anchor boxes such as calculating overlapping during training.
-
5 Sep 2017 85 repositories listed Syntology ran 2 of 9 samples · 7 unverified · 6 pointer-only (licence)Squeeze-and-Excitation Networks formed the foundation of our ILSVRC 2017 classification submission which won first place and reduced the top-5 error to 2.
-
9 Dec 2016 85 repositories listed Syntology ran 16 of 51 samples · 35 unverified · 11 pointer-only (licence)Feature pyramids are a basic component in recognition systems for detecting objects at different scales.
-
17 Sep 2014 83 repositories listed Syntology ran 27 of 42 samples · 15 unverified · 20 pointer-only (licence)We propose a deep convolutional neural network architecture codenamed "Inception", which was responsible for setting the new state of the art for classification and detection in the ImageNet Large-Scale Visual…
-
25 Mar 2021 80 repositories listed Syntology ran 108 of 207 samples · 99 unverified · 43 pointer-only (licence)This paper presents a new vision Transformer, called Swin Transformer, that capably serves as a general-purpose backbone for computer vision.
-
16 Apr 2019 76 repositories listed Syntology ran 10 of 130 samples · 120 unverifiedWe model an object as a single point --- the center point of its bounding box.
-
22 Nov 2017 68 repositories listed Syntology ran 2 of 4 samples · 2 unverified · 2 pointer-only (licence)In this work, we study 3D object detection from RGB-D data in both indoor and outdoor scenes.
-
6 May 2019 67 repositories listed Syntology ran 58 of 105 samples · 47 unverified · 46 pointer-only (licence)We achieve new state of the art results for mobile classification, detection and segmentation.
-
20 Nov 2019 64 repositories listed Syntology ran 11 of 70 samples · 59 unverified · 3 pointer-only (licence)Model efficiency has become increasingly important in computer vision.
-
24 Nov 2016 61 repositories listed Syntology ran 4 of 23 samples · 19 unverified · 4 pointer-only (licence)We present an approach to efficiently detect the 2D pose of multiple people in an image.
-
19 Jun 2017 59 repositories listed Syntology ran 9 of 17 samples · 8 unverified · 13 pointer-only (licence)Its principled nature also enables us to identify methods for both training and attacking neural networks that are reliable and, in a certain sense, universal.
-
11 Nov 2021 58 repositories listed Syntology ran 71 of 137 samples · 66 unverified · 73 pointer-only (licence)Our MAE approach is simple: we mask random patches of the input image and reconstruct the missing pixels.
-
10 Jan 2022 54 repositories listed Syntology ran 54 of 80 samples · 26 unverified · 11 pointer-only (licence)The "Roaring 20s" of visual recognition began with the introduction of Vision Transformers (ViTs), which quickly superseded ConvNets as the state-of-the-art image classification model.
-
10 Jan 2019 53 repositories listedCOCO test-dev results are up to 41.
-
20 May 2016 48 repositories listed Syntology ran 1 of 11 samples · 10 unverifiedIn contrast to previous region-based detectors such as Fast/Faster R-CNN that apply a costly per-region subnetwork hundreds of times, our region-based detector is fully convolutional with almost all computation shared…
-
7 Feb 2017 47 repositories listedWe adapted the join-training scheme of Faster RCNN framework from Caffe to TensorFlow as a baseline implementation for object detection.
-
17 Nov 2017 44 repositories listed Syntology ran 2 of 4 samples · 2 unverified · 4 pointer-only (licence)Accurate detection of objects in 3D point clouds is a central problem in many applications, such as autonomous navigation, housekeeping robots, and augmented/virtual reality.
Syntology lines on 27 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections