Browse State-of-the-Art › Scene Recognition
Scene Recognition
68 papers with code · 8 benchmarks · 16 datasets archive 2025-07-28
Benchmarks archive 2025-07-28
8 leaderboard tables shown for this task, 8 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
16 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
30 shown of 68 papers with code (207 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
27 Feb 2018 11 repositories listed Syntology ran 5 of 15 samples · 10 unverified · 5 pointer-only (licence)We demonstrate CSRNet on four datasets (ShanghaiTech dataset, the UCF_CC_50 dataset, the WorldEXPO'10 dataset, and the UCSD dataset) and we deliver the state-of-the-art performance.
-
6 Oct 2013 8 repositories listed Syntology ran 1 of 7 samples · 6 unverified · 1 pointer-only (licence)We evaluate whether features extracted from the activation of a deep convolutional network trained in a fully supervised fashion on a large, fixed set of object recognition tasks can be re-purposed to novel generic…
-
29 Apr 2015 4 repositories listedWe then present a systematic analysis of these networks and show that (1) the bilinear features are highly redundant and can be reduced by an order of magnitude in size without significant loss in accuracy, (2) are also…
-
23 Mar 2014 4 repositories listedWe report on a series of experiments conducted for different recognition tasks using the publicly available code and model of the \overfeat network which was trained to perform object classification on ILSVRC13.
-
18 May 2020 3 repositories listedIn this paper, we explore the problem of interesting scene prediction for mobile robots.
-
6 Apr 2022 2 repositories listedTo this end, we train different networks from scratch with the help of the largest RS scene recognition dataset up to now -- MillionAID, to obtain a series of RS pretrained backbones, including both convolutional neural…
-
20 Jan 2022 2 repositories listed Syntology ran 2 of 2 samples · 0 unverified · 1 pointer-only (licence)Prior work has studied different visual modalities in isolation and developed separate architectures for recognition of images, videos, and 3D data.
-
31 Aug 2020 2 repositories listed Syntology ran 3 of 5 samples · 2 unverified · 3 pointer-only (licence)Specifically, given an unlabeled video clip, we compute a series of spatio-temporal statistical summaries, such as the spatial location and dominant direction of the largest motion, the spatial location and dominant…
-
28 Feb 2020 2 repositories listedMoreover, we advocate multi-task learning as a way of improving scene recognition, building on the fact that the scene type is highly correlated with the objects in the scene, and therefore with its semantic…
-
10 Dec 2019 2 repositories listedThe hallucination task is treated as an auxiliary task, which can be used with any other action related task in a multitask learning setting.
-
3 Jan 2019 2 repositories listedSupervised machine learning based state-of-the-art computer vision techniques are in general data hungry.
-
4 Oct 2016 2 repositories listedConvolutional Neural Networks (CNNs) have made remarkable progress on scene recognition, partially due to these recent large-scale scene datasets, such as the Places and Places2.
-
7 Aug 2015 2 repositories listedWe verify the performance of trained Places205-VGGNet models on three datasets: MIT67, SUN397, and Places205.
-
9 Jan 2025 1 repository listedIn this study, we address this gap by constructing a large-scale ALS point cloud dataset and evaluating its impact on downstream applications.
-
1 Jan 2025 1 repository listedHowever, existing visual-to-visual and visual-to-textual Ego-Exo video alignment methods struggle with the problem that there could be non-visual overlap for the same activity.
-
4 Oct 2024 1 repository listed Syntology ran 5 of 9 samples · 4 unverified · 9 pointer-only (licence)To address this issue, we impose a requirement on learning systems to ensure that a new model strictly retains important capabilities of the old model while improving target-task performance, which we term model…
-
16 May 2024 1 repository listedContinuing with the above, we propose PIR-CLIP, a domain-specific CLIP-based framework for remote sensing image-text retrieval, to address semantic noise in remote sensing vision-language representations and further…
-
11 Dec 2023 1 repository listedVisual Question Answering (VQA) is one of the most important tasks in autonomous driving, which requires accurate recognition and complex situation evaluations.
-
9 Nov 2023 1 repository listedThis has been a significant bottleneck, particularly in the development of common sense reasoning and nuanced scene understanding necessary for safe and reliable autonomous driving.
-
4 Nov 2023 1 repository listedIn this paper, we propose a deep learning based crowd counting approach to automatically count number of manatees within a region, by using low quality images as input.
-
27 Oct 2023 1 repository listedOur highlight is the proposal of a paradigm that draws on prior knowledge to instruct adaptive learning of vision and text representations.
-
16 Jun 2023 1 repository listedIt consists of two stages, space granulation and attribute granulation.
-
23 May 2023 1 repository listedWe show that our questions 1) adequately represent the source material 2) can be used to diagnose a model's memory capacity 3) are not trivial for modern language models even when the memory demand does not exceed those…
-
15 May 2023 1 repository listedDespite the remarkable success of convolutional neural networks in various computer vision tasks, recognizing indoor scenes still presents a significant challenge due to their complex composition.
-
13 Mar 2023 1 repository listedMost deep learning backbones are evaluated on ImageNet.
-
13 Feb 2023 1 repository listedOur CoMAE presents a curriculum learning strategy to unify the two popular self-supervised representation learning algorithms: contrastive learning and masked image modeling.
-
15 Nov 2022 1 repository listed Syntology ran 0 of 1 samples · 1 unverifiedA shared goal of several machine learning communities like continual learning, meta-learning and transfer learning, is to design algorithms and models that efficiently and robustly adapt to unseen tasks.
-
20 Oct 2022 1 repository listed Syntology ran 2 of 11 samples · 9 unverifiedLongform media such as movies have complex narrative structures, with events spanning a rich variety of ambient visual scenes.
-
6 Sep 2022 1 repository listedCapsule networks are a neural network architecture specialized for visual scene recognition.
-
6 May 2022 1 repository listedFinally, our SSF allows our framework to learn the same scene scheme from multi-grain instance representations and fuses them, so that the entire framework is optimized as a whole.
Syntology lines on 7 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections