Browse State-of-the-Art › Object Counting
Object Counting
81 papers with code · 10 benchmarks · 28 datasets archive 2025-07-28
The goal of Object Counting task is to count the number of object instances in a single image or video sequence. It has many real-world applications such as traffic flow monitoring, crowdedness estimation, and product counting.
Source: Learning to Count Objects with Few Exemplar Annotations
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
10 leaderboard tables shown for this task, 10 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
28 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
4 subtasks in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 81 papers with code (158 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
25 Dec 2016 231 repositories listed Syntology ran 16 of 60 samples · 44 unverified · 22 pointer-only (licence)On the 156 classes not in COCO, YOLO9000 gets 16.
-
8 Jun 2015 144 repositories listed Syntology ran 80 of 148 samples · 68 unverified · 98 pointer-only (licence)A single neural network predicts bounding boxes and class probabilities directly from full images in one evaluation.
-
19 Jan 2023 7 repositories listed Syntology ran 5 of 14 samples · 9 unverified · 13 pointer-only (licence)This paper demonstrates an approach for learning highly semantic image representations without relying on hand-crafted data-augmentations.
-
14 Sep 2020 4 repositories listedProgress in the field of machine learning has been fueled by the introduction of benchmark datasets pushing the limits of existing algorithms.
-
17 May 2025 3 repositories listed Syntology ran 1 of 15 samples · 14 unverifiedLarge vision-language models exhibit inherent capabilities to handle diverse visual perception tasks.
-
14 Jul 2021 3 repositories listedIn terms of accuracy, YOLOv4-CSP was observed as the optimal model, with an AP@0.
-
28 Mar 2020 3 repositories listedThrough our analysis, we expect to make reasonable inference and prediction for the future development of crowd counting, and meanwhile, it can also provide feasible solutions for the problem of object counting in other…
-
7 Jan 2020 3 repositories listedVisual counting, a task that aims to estimate the number of objects from an image/video, is an open-set problem by nature, i.
-
25 Jul 2018 3 repositories listed Syntology ran 0 of 12 samples · 12 unverifiedHowever, we propose a detection-based method that does not need to estimate the size and shape of the objects and that outperforms regression-based methods.
-
3 Dec 2024 2 repositories listedThis comparative study contributes to the field of microbial image analysis by presenting innovative approaches to the recurring challenge of microorganism enumeration and by highlighting the capabilities of ViTs in the…
-
5 Jul 2024 2 repositories listed Syntology ran 4 of 9 samples · 5 unverifiedThe goal of this paper is to improve the generality and accuracy of open-vocabulary object counting in images.
-
20 May 2022 2 repositories listed Syntology ran 1 of 5 samples · 4 unverifiedSpecifically, we demonstrate that regression from vision transformer features without point-level supervision or reference images is superior to other reference-less methods and is competitive with methods that use…
-
8 Feb 2022 2 repositories listed Syntology ran 7 of 10 samples · 3 unverifiedIn this work, we investigate the visual reasoning capabilities and social biases of different text-to-image models, covering both multimodal transformer language models and diffusion models.
-
2 Sep 2020 2 repositories listedSupervised learning is often used to count objects in images, but for counting small, densely located objects, the required image annotations are burdensome to collect.
-
12 Mar 2020 2 repositories listedInspired by SFANet, the first model, which is named M-SFANet, is attached with atrous spatial pyramid pooling (ASPP) and context-aware module (CAN).
-
5 Mar 2020 2 repositories listedTo address this dilemma, we further propose an uncertainty-aware cross-modality vehicle detection (UA-CMDet) framework to extract complementary information from cross-modal images, which can significantly improve the…
-
6 Mar 2019 2 repositories listedMoreover, our approach improves state-of-the-art image-level supervised instance segmentation with a relative gain of 17.
-
14 Mar 2018 2 repositories listedAdding HR to a simple VGG front-end improves performance on all these benchmarks compared to a simple one-look baseline model and results in state-of-the-art performance for car counting.
-
1 Jan 2016 2 repositories listedEssentially, the CCNN is formulated as a regression model where the network learns how to map the appearance of the image patches to their corresponding object density maps.
-
11 Jul 2025 1 repository listedIn this paper, we investigate the applicability of the CLIP-EBC framework, originally designed for crowd counting, to car object counting using the CARPK dataset.
-
28 May 2025 1 repository listedTo support our approach, we analyze the key components of sota detection-based models and identify that detecting object centroids instead of bounding boxes is the key common factor behind their success in counting…
-
21 May 2025 1 repository listedWe believe the contributions of the proposed tasks, benchmark, and effective approach will advance future research in developing versatile object recognition systems.
-
10 Feb 2025 1 repository listedZero-shot counting is a subcategory of Generic Visual Object Counting, which aims to count objects from an arbitrary class in a given image.
-
1 Jan 2025 1 repository listedTo address this challenge, we propose a Hierarchical Semantic Correction Module that progressively refines text-image feature alignment, and a Representational Regional Coherence Loss that provides reliable supervision…
-
28 Nov 2024 1 repository listed Syntology ran 0 of 5 samples · 5 unverifiedTo address this gap in the geospatial domain, we present GEOBench-VLM, a comprehensive benchmark specifically designed to evaluate VLMs on geospatial tasks, including scene understanding, object counting, localization,…
-
27 Sep 2024 1 repository listedIn addition, a novel counting loss is proposed, that directly optimizes the detection task and avoids the issues of the standard surrogate loss.
-
24 Sep 2024 1 repository listedSpecifically, we argue that the current evaluation protocols do not measure the ability of the model to understand which object has to be counted.
-
26 Aug 2024 1 repository listedHowever, learning based on point annotations poses challenges due to the high imbalance between the sets of annotated and unannotated pixels, which is often treated with Gaussian smoothing of point annotations and focal…
-
6 Jul 2024 1 repository listedZero-shot object counting (ZOC) aims to enumerate objects in images using only the names of object classes during testing, without the need for manual annotations.
-
11 Jun 2024 1 repository listed Syntology ran 4 of 5 samples · 1 unverifiedThe unprecedented advancements in Multimodal Large Language Models (MLLMs) have demonstrated strong potential in interacting with humans through both language and visual inputs to perform downstream tasks such as visual…
Syntology lines on 10 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections