Browse State-of-the-Art › Scene Understanding
Scene Understanding
720 papers with code · 3 benchmarks · 43 datasets archive 2025-07-28
Scene understanding involves interpreting the visual information of a scene, including objects, their spatial relationships, and the overall layout. It goes beyond simple object recognition by considering the context and how objects relate to each other and the environment.
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
3 leaderboard tables shown for this task, 3 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| Semantic Scene Understanding Challenge (passive actuation & ground-truth localisation) (3 rows) | ACRV Baseline | — | — | — | Compare |
| ADE20K val (1 row) | CPN(ResNet-101) | Context Prior for Scene Segmentation | code | — | Compare |
| Semantic Scene Understanding Challenge (active actuation & ground-truth localisation) (1 row) | ACRV Baseline | — | — | — | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
43 datasets whose archive record lists this task, ordered by the archive's paper count. 30 shown of 43 until expanded.
Subtasks archive 2025-07-28
6 subtasks in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
30 shown of 720 papers with code (1,723 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
2 Nov 2015 74 repositories listed Syntology ran 9 of 44 samples · 35 unverified · 10 pointer-only (licence)We show that SegNet provides good performance with competitive inference time and more efficient inference memory-wise as compared to other architectures.
-
1 May 2014 38 repositories listed Syntology ran 1 of 7 samples · 6 unverifiedWe present a new dataset with the goal of advancing the state-of-the-art in object recognition by placing the question of object recognition in the context of the broader question of scene understanding.
-
26 Jul 2018 25 repositories listed Syntology ran 9 of 28 samples · 19 unverifiedIn this paper, we study a new task called Unified Perceptual Parsing, which requires the machine vision systems to recognize as many visual concepts as possible from a given image.
-
9 Nov 2015 21 repositories listed Syntology ran 0 of 18 samples · 18 unverifiedSemantic segmentation is an important tool for visual scene understanding and a meaningful measure of uncertainty is essential for decision making.
-
4 Jun 2018 15 repositories listed Syntology ran 17 of 24 samples · 7 unverified · 6 pointer-only (licence)Per-pixel ground-truth depth data is challenging to acquire at scale.
-
14 Jun 2017 14 repositories listed Syntology ran 3 of 4 samples · 1 unverified · 2 pointer-only (licence)As a result they are huge in terms of parameters and number of operations; hence slow too.
-
9 Oct 2017 13 repositories listedA comprehensive set of experiments on the publicly available Cityscapes dataset demonstrates that our system achieves an accuracy that is similar to the state of the art, while being orders of magnitude faster to…
-
17 Dec 2017 9 repositories listedAlthough CNN has shown strong capability to extract semantics from raw pixels, its capacity to capture spatial relationships of pixels across rows and columns of an image is not fully explored.
-
1 Nov 2019 8 repositories listedWe leverage this scaling to train an agent for 2.
-
5 Dec 2020 7 repositories listedThis dataset demonstrates the post flooded damages of the affected areas.
-
1 Apr 2019 7 repositories listed Syntology ran 1 of 13 samples · 12 unverifiedScene understanding of high resolution aerial images is of great importance for the task of automated monitoring in various remote sensing applications.
-
10 Oct 2018 7 repositories listed Syntology ran 2 of 19 samples · 17 unverifiedThese algorithms are not directly applicable to large-scale learning problems since they scale poorly with the dimensionality of the gradients and the number of tasks.
-
22 Feb 2022 6 repositories listedWith only text supervision and without any pixel-level annotations, GroupViT learns to group together semantic regions and successfully transfers to the task of semantic segmentation in a zero-shot manner, i.
-
8 Jul 2019 6 repositories listed Syntology ran 1 of 8 samples · 7 unverified3D object detection from LiDAR point cloud is a challenging problem in 3D scene understanding and has many practical applications.
-
27 Nov 2018 6 repositories listedCompared with real-time segmentation models such as BiSeNet, our model achieves higher accuracy at comparable speed on the Cityscapes Dataset, enabling the application in speed-demanding tasks such as street-scene…
-
27 Feb 2017 6 repositories listed Syntology ran 1 of 1 samples · 0 unverifiedWe introduce a method for training GANs with discrete data that uses the estimated difference measure from the discriminator to compute importance weights for generated samples, thus providing a policy gradient for…
-
20 Nov 2019 5 repositories listedWith the help of novel masks or scenes, we enhance the current datasets using synthesized shadow images.
-
2 Apr 2019 5 repositories listed Syntology ran 1 of 8 samples · 7 unverifiedDespite the relevance of semantic scene understanding for this application, there is a lack of a large dataset for this task which is based on an automotive LiDAR.
-
6 Aug 2016 5 repositories listedTo remedy this, we develop a method for adapting deep features to align with human similarity judgments, resulting in image representations that can potentially be used to extend the scope of psychological experiments.
-
4 Dec 2023 4 repositories listed Syntology ran 14 of 26 samples · 12 unverifiedMonocular depth estimation is a fundamental computer vision task.
-
24 Sep 2023 4 repositories listed Syntology ran 0 of 12 samples · 12 unverifiedAs the application scenarios of mobile robots are getting more complex and challenging, scene understanding becomes increasingly crucial.
-
22 Jun 2021 4 repositories listedA popular solution to this problem is to use a single pooling operation to reduce the sequence length.
-
30 Mar 2021 4 repositories listed Syntology ran 9 of 13 samples · 4 unverified · 13 pointer-only (licence)Understanding the scene around the ego-vehicle is key to assisted and autonomous driving.
-
24 Mar 2021 4 repositories listedMoreover, we also propose a new head detector, HeadHunter, which is designed for small head detection in crowded scenes.
-
16 Apr 2020 4 repositories listed Syntology ran 2 of 6 samples · 4 unverifiedIn this technical report, we present two novel datasets for image scene understanding.
-
3 Apr 2020 4 repositories listed Syntology ran 0 of 5 samples · 5 unverifiedInstance segmentation is an important task for scene understanding.
-
7 Jan 2020 4 repositories listedFew works have studied the disambiguating contribution of subsidiary relations made available via graph networks.
-
17 Nov 2019 4 repositories listedBeing natural, touchless, and fun-embracing, language-based inputs have been demonstrated effective for various tasks from image generation to literacy education for children.
-
13 Jul 2019 4 repositories listed Syntology ran 1 of 21 samples · 20 unverifiedUnderstanding the world in 3D is a critical component of urban autonomous driving.
-
2 May 2018 4 repositories listed Syntology ran 0 of 4 samples · 4 unverifiedIt is valuable to fuse outputs from multiple sensors to boost overall performance.
Syntology lines on 18 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections