Browse State-of-the-Art › Weakly-Supervised Object Localization
Weakly-Supervised Object Localization
82 papers with code · 8 benchmarks · 3 datasets archive 2025-07-28
Weakly supervised object localization (WSOL) learns to localize objects with only image-level labels, no object level labels (bonding boxes, etc.,) is needed. It is more attractive since image-level labels are much easier and cheaper to obtain.
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
8 leaderboard tables shown for this task, 8 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
3 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
30 shown of 82 papers with code (140 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
14 Dec 2015 35 repositories listed Syntology ran 2 of 17 samples · 15 unverified · 3 pointer-only (licence)In this work, we revisit the global average pooling layer proposed in [13], and shed light on how it explicitly enables the convolutional neural network to have remarkable localization ability despite being trained on…
-
22 Jun 2021 3 repositories listedTo evaluate the quality of the class activation maps produced by LayerCAM, we apply them to weakly-supervised object localization and semantic segmentation.
-
1 Aug 2020 3 repositories listedAt the heart of this progress is convolutional neural networks (CNNs) that are capable of learning representations or features given a set of data.
-
13 Apr 2017 3 repositories listedWe propose `Hide-and-Seek', a weakly-supervised framework that aims to improve object localization in images and action localization in videos.
-
22 Sep 2023 2 repositories listedIn addition, our method also achieves state-of-the-art weakly supervised semantic segmentation performance on the PASCAL VOC 2012 and MS COCO 2014 datasets.
-
21 Jul 2022 2 repositories listed Syntology ran 7 of 10 samples · 3 unverifiedWeakly Supervised Object Localization (WSOL), which aims to localize objects by only using image-level labels, has attracted much attention because of its low annotation cost in real applications.
-
25 Mar 2022 2 repositories listed Syntology ran 2 of 2 samples · 0 unverifiedWhile class activation map (CAM) generated by image classification network has been widely used for weakly supervised object localization (WSOL) and semantic segmentation (WSSS), such classifiers usually focus on…
-
1 Dec 2021 2 repositories listedExisting FPM-based methods use cross-entropy (CE) to evaluate the foreground prediction map and to guide the learning of generator.
-
29 Sep 2021 2 repositories listedWe also show that training a class-agnostic detector on the discovered objects boosts results by another 7 points.
-
27 Mar 2021 2 repositories listedTS-CAM finally couples the patch tokens with the semantic-agnostic attention map to achieve semantic-aware localization.
-
19 Jul 2020 2 repositories listedCompared to classification networks, attention visualization for retrieval networks is hardly studied.
-
8 Jul 2020 2 repositories listedIn this paper, we argue that WSOL task is ill-posed with only image-level labels, and propose a new evaluation protocol where full supervision is limited to only a small held-out set not overlapping with the test set.
-
21 Jan 2020 2 repositories listed Syntology ran 2 of 2 samples · 0 unverified · 1 pointer-only (licence)In this paper, we argue that WSOL task is ill-posed with only image-level labels, and propose a new evaluation protocol where full supervision is limited to only a small held-out set not overlapping with the test set.
-
19 Apr 2018 2 repositories listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)With such an adversarial learning, the two parallel-classifiers are forced to leverage complementary object regions for classification and can finally generate integral object localization together.
-
1 Jul 2017 2 repositories listedThis paper introduces WILDCAT, a deep learning method which jointly aims at aligning image regions for gaining spatial invariance and learning strongly localized features.
-
31 Mar 2025 1 repository listedThey also face the well-known issue of asynchronous convergence between classification and localization tasks.
-
22 Jan 2025 1 repository listedWeakly supervised object localization (WSOL) using classification models trained with only image-class labels remains an important challenge in computer vision.
-
8 Jul 2024 1 repository listedOur TrCAM-V method allows training a localization network by sampling pseudo-pixels on the fly from these regions.
-
23 May 2024 1 repository listedA fundamental limitation to the development of 3D saliency methods is the lack of a benchmark to quantitatively assess these on 3D data.
-
29 Apr 2024 1 repository listedA WSOL model initially trained on some labeled source image data can be adapted using unlabeled target data in cases of significant domain shifts caused by variations in staining, scanners, and cancer type.
-
15 Apr 2024 1 repository listedThese bboxes are also employed to estimate the threshold from LOC maps, circumventing the need for test-set bbox annotations.
-
11 Mar 2024 1 repository listedThe reason for the high-performance of large kernel CNNs in downstream tasks has been attributed to the large effective receptive field (ERF) produced by large size kernels, but this view has not been fully tested.
-
16 Dec 2023 1 repository listedWhile foveation enables it to process different regions of the input with variable degrees of detail, saccades allow it to change the focus point of such foveated regions.
-
9 Oct 2023 1 repository listedSubsequently, these proposals are used as pseudo-labels to train our new transformer-based WSOL model designed to perform classification and localization tasks.
-
17 Sep 2023 1 repository listedTo the best of our knowledge, we are the first to address this task.
-
19 Jul 2023 1 repository listed Syntology ran 1 of 1 samples · 0 unverifiedDuring training, GenPromp converts image category labels to learnable prompt embeddings which are fed to a generative model to conditionally recover the input image with noise and learn representative embeddings.
-
17 Apr 2023 1 repository listedTo handle such data, we propose a novel paradigm of contrastive representation co-learning using both labeled and unlabeled data to generate a complete G-CAM (Generalized Class Activation Map) for object localization,…
-
18 Mar 2023 1 repository listed Syntology ran 5 of 6 samples · 1 unverified · 6 pointer-only (licence)Specifically, a spatial token is first introduced in the input space to aggregate representations for localization task.
-
3 Jan 2023 1 repository listedPrevious weakly-supervised object localization (WSOL) methods aim to expand activation map discriminative areas to cover the whole objects, yet neglect two inherent challenges when relying solely on image-level labels.
-
1 Jan 2023 1 repository listedVision transformers use [CLS] tokens to predict image classes.
Syntology lines on 7 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections