Browse State-of-the-Art › Object Localization
Object Localization
282 papers with code · 18 benchmarks · 18 datasets archive 2025-07-28
Object Localization is the task of locating an instance of a particular object category in an image, typically by specifying a tightly cropped bounding box centered on the instance. An object proposal specifies a candidate bounding box, and an object proposal is said to be a correct localization if it sufficiently overlaps a human-labeled “ground-truth” bounding box for the given object. In the literature, the “Object Localization” task is to locate one instance of an object category, whereas “object detection” focuses on locating all instances of a category in a given image.
Source: Fast On-Line Kernel Density Estimation for Active Object Localization
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
18 leaderboard tables shown for this task, 18 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted. 10 shown of 18 until expanded.
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
18 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
5 subtasks in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 282 papers with code (617 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
20 Mar 2017 179 repositories listed Syntology ran 42 of 140 samples · 98 unverified · 23 pointer-only (licence)Our approach efficiently detects objects in an image while simultaneously generating a high-quality segmentation mask for each instance.
-
22 Nov 2017 68 repositories listed Syntology ran 2 of 4 samples · 2 unverified · 2 pointer-only (licence)In this work, we study 3D object detection from RGB-D data in both indoor and outdoor scenes.
-
17 Nov 2017 44 repositories listed Syntology ran 2 of 4 samples · 2 unverified · 4 pointer-only (licence)Accurate detection of objects in 3D point clouds is a central problem in many applications, such as autonomous navigation, housekeeping robots, and augmented/virtual reality.
-
1 May 2014 38 repositories listed Syntology ran 1 of 7 samples · 6 unverifiedWe present a new dataset with the goal of advancing the state-of-the-art in object recognition by placing the question of object recognition in the context of the broader question of scene understanding.
-
14 Dec 2015 35 repositories listed Syntology ran 2 of 17 samples · 15 unverified · 3 pointer-only (licence)In this work, we revisit the global average pooling layer proposed in [13], and shed light on how it explicitly enables the convolutional neural network to have remarkable localization ability despite being trained on…
-
13 May 2019 30 repositories listed Syntology ran 17 of 24 samples · 7 unverified · 5 pointer-only (licence)Regional dropout strategies have been proposed to enhance the performance of convolutional neural network classifiers.
-
30 Oct 2017 24 repositories listed Syntology ran 4 of 7 samples · 3 unverified · 3 pointer-only (licence)Over the last decade, Convolutional Neural Network (CNN) models have been highly successful in solving complex vision problems.
-
20 Mar 2017 7 repositories listed Syntology ran 0 of 6 samples · 6 unverifiedBridging the 'reality gap' that separates simulated robotics from experiments on hardware could accelerate robotic research through improved data availability.
-
15 Aug 2021 6 repositories listed Syntology ran 0 of 4 samples · 4 unverifiedIn this paper, we identify that the problem is that the binary classifiers in existing proposal methods tend to overfit to the training categories.
-
20 Jun 2018 6 repositories listedIn these networks, the training procedure usually requires providing bounding boxes or the maximum number of expected objects.
-
15 Sep 2020 5 repositories listed Syntology ran 4 of 30 samples · 26 unverifiedThis paper presents the evaluation methodology, datasets, and results of the BOP Challenge 2020, the third in a series of public competitions organized with the goal to capture the status quo in the field of 6D object…
-
19 Feb 2025 4 repositories listed Syntology ran 2 of 3 samples · 1 unverifiedWe introduce Qwen2.
-
3 Aug 2019 4 repositories listed Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)With the intention to create an enhanced visual explanation in terms of visual sharpness, object localization and explaining multiple occurrences of objects in a single image, we present Smooth Grad-CAM++…
-
23 Sep 2018 4 repositories listed Syntology ran 0 of 4 samples · 4 unverifiedLarge-scale object detection datasets (e.
-
6 Mar 2018 4 repositories listedThe proposed approach significantly improves the state-of-the-art for monocular object localization on arbitrarily-shaped roads.
-
12 Feb 2024 3 repositories listed Syntology ran 1 of 5 samples · 4 unverified · 4 pointer-only (licence)Visually-conditioned language models (VLMs) have seen growing adoption in applications such as visual dialogue, scene understanding, and robotic task planning; adoption that has fueled a wealth of new models such as…
-
17 Dec 2021 3 repositories listedIn this paper, we comprehensively study three architecture design choices on ViT -- spatial reduction, doubled channels, and multiscale features -- and demonstrate that a vanilla ViT architecture can fulfill this goal…
-
29 Nov 2021 3 repositories listedA typical pipeline for multi-object tracking (MOT) is to use a detector for object localization, and following re-identification (re-ID) for object association.
-
22 Jun 2021 3 repositories listedTo evaluate the quality of the class activation maps produced by LayerCAM, we apply them to weakly-supervised object localization and semantic segmentation.
-
1 Aug 2020 3 repositories listedAt the heart of this progress is convolutional neural networks (CNNs) that are capable of learning representations or features given a set of data.
-
2 Jul 2020 3 repositories listedWe evaluated our model on several medical (ACDC, LVSC, CHAOS) and non-medical (PPSS) datasets, and we report performance levels matching those achieved by models trained with fully annotated segmentation masks.
-
18 Dec 2019 3 repositories listedWe introduce the task of 3D object localization in RGB-D scans using natural language descriptions.
-
28 May 2017 3 repositories listedConvolutional networks for image classification progressively reduce resolution until the image is represented by tiny feature maps in which the spatial structure of the scene is no longer discernible.
-
13 Apr 2017 3 repositories listedWe propose `Hide-and-Seek', a weakly-supervised framework that aims to improve object localization in images and action localization in videos.
-
18 Nov 2015 3 repositories listedWe present an active detection model for localizing objects in scenes.
-
28 Apr 2024 2 repositories listed Syntology ran 11 of 11 samples · 0 unverified · 11 pointer-only (licence)Specifically, our Mamba-based tracker achieves 43.
-
30 Jan 2024 2 repositories listedCPR reduces the semantic variance by selecting a semantic centre point in a neighbourhood region to replace the initial annotated point.
-
22 Sep 2023 2 repositories listedIn addition, our method also achieves state-of-the-art weakly supervised semantic segmentation performance on the PASCAL VOC 2012 and MS COCO 2014 datasets.
-
25 Nov 2022 2 repositories listedIn recent works on semantic segmentation, there has been a significant focus on designing and integrating transformer-based encoders.
-
20 Nov 2022 2 repositories listedIn this paper, we propose a single-stage backbone network for Color-Event Unified Tracking (CEUTrack), which achieves the above functions simultaneously.
Syntology lines on 15 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections